2022 benchmark; English release tracked separately
DDXPlus ↗
The most likely disease and the useful differential are different targets.
Diagnostic evaluation studies
Independent analysis of DDXPlus and MedR-Bench: synthetic differentials, case-report reasoning, examination requests, oracle conditions, metric denominators and published results.
Start with the supplied initial evidence.
Ask about symptoms or antecedents.
Return an underlying pathology or differential distribution.
Use one-turn, free-turn or oracle access to examination evidence.
Request examinations or produce a diagnosis/treatment plan.
Match outputs to the versioned references with the source judge.
Simulator-derived differentials and case-report judgments have different reference standards. Follow each dossier's information setting and metric.
Top-pathology accuracy and differential coverage answer different questions.
Requested examinations and oracle evidence produce different tasks.
Written-rationale scores use separate denominators and an automated judge.
DDXPlus separates a simulated pathology from its reference differential. MedR-Bench separates examination gathering, diagnosis, treatment and written reasoning. We analyze those task boundaries and the original scoring rules, with selected version-matched historical results. Our coverage explorer and guides are original analytical resources. The benchmarks and experiments belong to their credited authors; this site has not run the displayed models or created a new patient dataset.
Read the original evidence closely
2022 benchmark; English release tracked separately
The most likely disease and the useful differential are different targets.
Nature Communications 2025, seven-model study
Supplying the examinations changes what a diagnosis score means.
An original analytical tool
Compare actual simulator and case-report tasks. Filter by diagnostic target, information gathering, clinical stage or written-reasoning evaluation.
9 of 9 evidence entries shown
Case-report treatment is the reference; not a prospective intervention comparison.
[4]This is our independent analytical map of published tasks. It does not execute an evaluation, predict a model’s performance or establish clinical benefit. [1][2][3][4][5][6]
Original analysis / methods and interpretation
Interpret differential recall, precision and pathology accuracy as different diagnostic targets.
Understand what changes between one-turn, free-turn and oracle diagnosis conditions.
Analyze efficiency, factuality and completeness without treating written explanations as hidden model reasoning.
Questions, answered
Specific tasks. Stated conditions.
Inspect every source.
No. It contains synthetic patients generated from a knowledge base and rule-based diagnostic system. Its diagnostic scores measure agreement with that constructed reference.
The model receives all ground-truth examination evidence before diagnosis. This differs from identifying which examinations to request.
They evaluate the written reasoning that is supplied to the evaluator. They do not directly establish which internal computation caused the answer.
The November 2025 Nature Communications study and its supplementary tables, rather than mixing those results with the earlier five-model preprint.
Working tool / saved on this device
Use this secondary checklist to document a run or literature comparison after inspecting the named benchmark conditions. Completion records documentation, not performance.
Evidence you can inspect. Benchmark dossiers distinguish published facts from our interpretation, with source versions and access notes attached.
Download the evidence ↗