Independent benchmark analysis / Nature Communications 2025, seven-model study
MedR-Bench
Supplying the examinations changes what a diagnosis score means.
MedR-Bench reorganizes published case reports into staged clinical tasks. It evaluates examination recommendations, diagnoses and treatment plans, while a separate automated evaluator scores written reasoning. Our analysis follows the information made available at each stage: one-turn diagnosis and oracle diagnosis are deliberately different conditions. We reproduce selected version-matched measurements and explain why factual reasoning, complete reasoning and a correct final answer must remain distinct. The source cases are real published reports, but the interactions and scoring are reconstructed research procedures rather than a prospective clinical study. This site contributes interpretation, not new experimental results.
01 / What is being tested?
The task, before the score.
- input
- Structured case information, with examinations withheld or supplied according to setting
- output
- Requested examinations, diagnosis or treatment plan plus written rationale
- unit
- One case-report-derived patient case
- setting
- One-turn, free-turn and oracle conditions; reasoning judged through an agentic GPT-4o-based pipeline
Data origin. PMC Open Access case reports reorganized into structured cases using GPT-4o. [4][5][6]
- Structured cases
- 1,453
957 diagnosis cases plus 496 treatment cases.
Patient cases [4] - Rare-disease cases
- 656
491 diagnosis and 165 treatment cases; these are subsets, not extra cases.
Patient cases [4] - Body systems
- 13
Source taxonomy coverage.
Abstract [4] - Clinical stages
- 3
Examination recommendation, diagnosis and treatment.
Evaluation settings [4] - Study version
- 7 models
The journal study expands the five-model March 2025 preprint. Use version-matched results.
Abstract; evaluated models [4]
- 01
Choose information availability
Use one-turn, free-turn or oracle access to examination evidence.
[4] - 02
Generate clinical outputs
Request examinations or produce a diagnosis/treatment plan.
[6] - 03
Evaluate outcomes
Match outputs to the versioned references with the source judge.
[4] - 04
Evaluate written reasoning
Decompose steps, verify factuality and match reference steps.
[4]
Dataset anatomy
Case types, without double-counting rare diseases
Contains 491 rare-disease cases.
Contains 165 rare-disease cases.
Rare-disease labels overlap these task subsets. They must not be added as a third disjoint group. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Outcome accuracy and written-reasoning dimensions
Higher is better within the same condition
Outcome matching uses GPT-4o; efficiency, factuality and completeness use distinct step denominators. Completeness is unavailable for examination recommendation because the case reports rarely document why examinations were selected.
Efficiency = effective / generated steps; factuality = correct effective / effective steps; completeness = covered reference / reference steps
Written rationale evaluation is not direct observation of a model’s internal causal reasoning. Freeze judge, retrieval and prompts. [4][6]
03 / Measured evidence
Results, with their conditions attached.
Paper-reported results / selected rows
Diagnosis after one examination turn
2025 journal supplement, all 957 diagnosis cases; one-turn examination condition.
Paper-reported values, not our runs. Source uses two-sided 95% intervals. Generation and judge configurations are documented in Supplement §A.2.
Source: Supplementary Table 3 [5]
Paper-reported results / selected rows
Diagnosis with all reference examinations
Same journal study and diagnosis cases; oracle evidence condition.
The evidence condition changes. This panel is not a matched deployment comparison or a new-model release trend.
Source: Supplementary Table 3 [5]
04 / Our original analysis
What follows from the design?
Reasoning dimensions have different denominators
Factuality considers effective steps; completeness considers reference steps. [4]
Our inference: a short entirely factual rationale can still omit most required reasoning, so the metrics should not be collapsed.
Case reports create a publication filter
Structured cases originate in published case reports. [4]
Our inference: their disease mix and narrative completeness need not represent consecutive patients in a clinic.
A judge is part of the measurement
05 / Scope of the evidence
Where this benchmark stops.
Reconstructed patient interactions
The benchmark uses staged case information and an agentic patient role, not a prospective clinical encounter. [4]
Written reasoning reference
A published discussion cannot document every valid reasoning route or expose hidden model computation. [4]
Automated judgment dependency
Semantic matching, web retrieval and judge behavior can affect the scores. [4]
- Availability
- Public repository and supplements
- License
- Repository identifies CC BY-SA; version not specified in its short LICENSE
- Conditions
- Consult the creators for exact reuse terms and original article rights before redistributing cases.
Evidence trail
Read the originals.
- Quantifying the reasoning abilities of LLMs on clinical cases ↗
Qiu et al. / Nature Communications. Peer-reviewed seven-model study, case counts, written-reasoning metrics and judge implementation.
- MedR-Bench Supplementary Information ↗
Qiu et al. / Nature Communications. Version-matched results with confidence intervals and generation parameters.
- MedR-Bench official code and data ↗
MAGIC-AI4Med. Source code, dataset and saved model responses; repository license identifies CC BY-SA without a version.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.