Independent benchmark analysis / Nature Communications 2025, seven-model study

MedR-Bench

Supplying the examinations changes what a diagnosis score means.

MedR-Bench reorganizes published case reports into staged clinical tasks. It evaluates examination recommendations, diagnoses and treatment plans, while a separate automated evaluator scores written reasoning. Our analysis follows the information made available at each stage: one-turn diagnosis and oracle diagnosis are deliberately different conditions. We reproduce selected version-matched measurements and explain why factual reasoning, complete reasoning and a correct final answer must remain distinct. The source cases are real published reports, but the interactions and scoring are reconstructed research procedures rather than a prospective clinical study. This site contributes interpretation, not new experimental results.

01 / What is being tested?

The task, before the score.

input
Structured case information, with examinations withheld or supplied according to setting
output
Requested examinations, diagnosis or treatment plan plus written rationale
unit
One case-report-derived patient case
setting
One-turn, free-turn and oracle conditions; reasoning judged through an agentic GPT-4o-based pipeline

Data origin. PMC Open Access case reports reorganized into structured cases using GPT-4o. [4][5][6]

Structured cases
1,453

957 diagnosis cases plus 496 treatment cases.

Patient cases [4]
Rare-disease cases
656

491 diagnosis and 165 treatment cases; these are subsets, not extra cases.

Patient cases [4]
Body systems
13

Source taxonomy coverage.

Abstract [4]
Clinical stages
3

Examination recommendation, diagnosis and treatment.

Evaluation settings [4]
Study version
7 models

The journal study expands the five-model March 2025 preprint. Use version-matched results.

Abstract; evaluated models [4]
  1. 01

    Choose information availability

    Use one-turn, free-turn or oracle access to examination evidence.

    [4]
  2. 02

    Generate clinical outputs

    Request examinations or produce a diagnosis/treatment plan.

    [6]
  3. 03

    Evaluate outcomes

    Match outputs to the versioned references with the source judge.

    [4]
  4. 04

    Evaluate written reasoning

    Decompose steps, verify factuality and match reference steps.

    [4]

Dataset anatomy

Case types, without double-counting rare diseases

Diagnosis

Contains 491 rare-disease cases.

957 cases[4]
Treatment

Contains 165 rare-disease cases.

496 cases[4]

Rare-disease labels overlap these task subsets. They must not be added as a third disjoint group. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Outcome accuracy and written-reasoning dimensions

Higher is better within the same condition

Outcome matching uses GPT-4o; efficiency, factuality and completeness use distinct step denominators. Completeness is unavailable for examination recommendation because the case reports rarely document why examinations were selected.

Scoring definition

Efficiency = effective / generated steps; factuality = correct effective / effective steps; completeness = covered reference / reference steps

Written rationale evaluation is not direct observation of a model’s internal causal reasoning. Freeze judge, retrieval and prompts. [4][6]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Diagnosis after one examination turn

2025 journal supplement, all 957 diagnosis cases; one-turn examination condition.

Accuracy · %
050100
Reported
OpenAI-o3-mini95% CI 61.97–68.02; o3-mini-2025-01-31.
64.99%
Gemini-2.0-FT95% CI 65.60–71.49; flash-thinking-exp-01-21.
68.55%
DeepSeek-R195% CI 68.93–74.64; 671B model.
71.79%

Paper-reported values, not our runs. Source uses two-sided 95% intervals. Generation and judge configurations are documented in Supplement §A.2.

Source: Supplementary Table 3 [5]

Paper-reported results / selected rows

Diagnosis with all reference examinations

Same journal study and diagnosis cases; oracle evidence condition.

Accuracy · %
050100
Reported
OpenAI-o3-mini95% CI 81.58–86.24.
83.91%
Gemini-2.0-FT95% CI 84.69–88.98.
86.83%
DeepSeek-R195% CI 87.84–91.68.
89.76%

The evidence condition changes. This panel is not a matched deployment comparison or a new-model release trend.

Source: Supplementary Table 3 [5]

04 / Our original analysis

What follows from the design?

01

Information gathering is separately measurable

Published evidence

The same benchmark includes one-turn and oracle diagnosis conditions. [4][5]

Our interpretation

Our inference: report both before treating a high oracle score as evidence that a system knows which examinations to request.

02

Reasoning dimensions have different denominators

Published evidence

Factuality considers effective steps; completeness considers reference steps. [4]

Our interpretation

Our inference: a short entirely factual rationale can still omit most required reasoning, so the metrics should not be collapsed.

03

Case reports create a publication filter

Published evidence

Structured cases originate in published case reports. [4]

Our interpretation

Our inference: their disease mix and narrative completeness need not represent consecutive patients in a clinic.

04

A judge is part of the measurement

Published evidence

The evaluator uses language models and external search. [4][6]

Our interpretation

Our inference: preserve retrieval snapshots and judge configuration when a result needs to be audited later.

05 / Scope of the evidence

Where this benchmark stops.

Reconstructed patient interactions

The benchmark uses staged case information and an agentic patient role, not a prospective clinical encounter. [4]

Written reasoning reference

A published discussion cannot document every valid reasoning route or expose hidden model computation. [4]

Automated judgment dependency

Semantic matching, web retrieval and judge behavior can affect the scores. [4]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Public repository and supplements
License
Repository identifies CC BY-SA; version not specified in its short LICENSE
Conditions
Consult the creators for exact reuse terms and original article rights before redistributing cases.
[6]

Evidence trail

Read the originals.

  1. Quantifying the reasoning abilities of LLMs on clinical cases ↗

    Qiu et al. / Nature Communications. Peer-reviewed seven-model study, case counts, written-reasoning metrics and judge implementation.

  2. MedR-Bench Supplementary Information ↗

    Qiu et al. / Nature Communications. Version-matched results with confidence intervals and generation parameters.

  3. MedR-Bench official code and data ↗

    MAGIC-AI4Med. Source code, dataset and saved model responses; repository license identifies CC BY-SA without a version.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗