Diagnostic evaluation studies

Follow the evidence into the diagnosis.

Independent analysis of DDXPlus and MedR-Bench: synthetic differentials, case-report reasoning, examination requests, oracle conditions, metric denominators and published results.

Diagnostic reasoning / evidence-to-output map

The path to a conclusion is part of the task

01

DDXPlus

Synthetic patient interactions

  1. Initialize a synthetic patient

    Start with the supplied initial evidence.

  2. Collect structured evidence

    Ask about symptoms or antecedents.

  3. Stop and predict

    Return an underlying pathology or differential distribution.

Approximately 1.3 million · synthetic population [1][2][3]

02

MedR-Bench

Case-report-derived patient cases

  1. Choose information availability

    Use one-turn, free-turn or oracle access to examination evidence.

  2. Generate clinical outputs

    Request examinations or produce a diagnosis/treatment plan.

  3. Evaluate outcomes

    Match outputs to the versioned references with the source judge.

1,453 · structured cases [4][5][6]

Simulator-derived differentials and case-report judgments have different reference standards. Follow each dossier's information setting and metric.

01 /

Diagnostic target

Top-pathology accuracy and differential coverage answer different questions.

02 /

Information availability

Requested examinations and oracle evidence produce different tasks.

03 /

Reasoning evidence

Written-rationale scores use separate denominators and an automated judge.

Our analytical question

DDXPlus separates a simulated pathology from its reference differential. MedR-Bench separates examination gathering, diagnosis, treatment and written reasoning. We analyze those task boundaries and the original scoring rules, with selected version-matched historical results. Our coverage explorer and guides are original analytical resources. The benchmarks and experiments belong to their credited authors; this site has not run the displayed models or created a new patient dataset.

Read the original evidence closely

The benchmark, unpacked.

Primary source library ↗
Dossier01

2022 benchmark; English release tracked separately

DDXPlus ↗

The most likely disease and the useful differential are different targets.

UnitOne synthetic patient interactionMeasureDifferential recall, precision, F1 and pathology accuracy
Dossier02

Nature Communications 2025, seven-model study

MedR-Bench ↗

Supplying the examinations changes what a diagnosis score means.

UnitOne case-report-derived patient caseMeasureOutcome accuracy and written-reasoning dimensions

An original analytical tool

Map the diagnostic target and information path

Evidence explorer

Compare actual simulator and case-report tasks. Filter by diagnostic target, information gathering, clinical stage or written-reasoning evaluation.

9 of 9 evidence entries shown

Diagnostic target

DDXPlus: top pathology

Read dossier ↗
Input
Collected simulator evidence
Output
Most likely pathology
Measure
GTPA@1
Interpretation boundary

Synthetic underlying pathology is the target.

[1]
Diagnostic target

DDXPlus: differential coverage

Read dossier ↗
Input
Collected simulator evidence
Output
Set of plausible pathologies
Measure
DDR / DDP / DDF1
Interpretation boundary

Apply the published probability threshold before set comparison.

[1]
Information gathering

DDXPlus: evidence gathering

Read dossier ↗
Input
Initial symptom and interactive evidence
Output
Evidence requests
Measure
Positive evidence recall; interaction length
Interpretation boundary

A longer interaction is not automatically better.

[1]
Information gathering

MedR: one-turn examinations

Read dossier ↗
Input
Case without ancillary test findings
Output
Recommended examination list
Measure
Precision and recall
Interpretation boundary

References describe examinations present in the source case report.

[4]
Information gathering

MedR: free-turn examinations

Read dossier ↗
Input
Case with iterative examination access
Output
Sequence of requests
Measure
Precision and recall
Interpretation boundary

The model chooses when to stop; patient responses are agent-mediated.

[4]
Clinical stage

MedR: one-turn diagnosis

Read dossier ↗
Input
Case plus once-requested examinations
Output
Diagnosis
Measure
Semantic accuracy
Interpretation boundary

Information availability depends on the preceding request.

[4]
Clinical stage

MedR: oracle diagnosis

Read dossier ↗
Input
All reference examinations
Output
Diagnosis
Measure
Semantic accuracy
Interpretation boundary

Does not test selecting the examinations.

[4]
Clinical stage

MedR: oracle treatment

Read dossier ↗
Input
Complete case including diagnosis
Output
Treatment plan
Measure
Semantic plan accuracy
Interpretation boundary

Case-report treatment is the reference; not a prospective intervention comparison.

[4]
Reasoning

MedR: written-reasoning evaluation

Read dossier ↗
Input
Generated rationale and reference reasoning
Output
Judged steps
Measure
Efficiency, factuality, completeness
Interpretation boundary

Measures expressed rationale with an automated evaluator. Examination recommendation lacks reference reasoning for completeness scoring.

[4][6]

This is our independent analytical map of published tasks. It does not execute an evaluation, predict a model’s performance or establish clinical benefit. [1][2][3][4][5][6]

Original analysis / methods and interpretation

What the score leaves unsaid.

All analyses →

Questions, answered

Read the result in context.

Specific tasks. Stated conditions.
Inspect every source.

Does DDXPlus contain real patient records?

No. It contains synthetic patients generated from a knowledge base and rule-based diagnostic system. Its diagnostic scores measure agreement with that constructed reference.

What is the MedR-Bench oracle condition?

The model receives all ground-truth examination evidence before diagnosis. This differs from identifying which examinations to request.

Do reasoning scores reveal a model’s internal thinking?

They evaluate the written reasoning that is supplied to the evaluator. They do not directly establish which internal computation caused the answer.

Which MedR-Bench version do the result panels use?

The November 2025 Nature Communications study and its supplementary tables, rather than mixing those results with the earlier five-model preprint.

Working tool / saved on this device

Prepare a benchmark comparison brief

Interactive worksheet

Use this secondary checklist to document a run or literature comparison after inspecting the named benchmark conditions. Completion records documentation, not performance.

Define the diagnostic target

Evidence you can inspect. Benchmark dossiers distinguish published facts from our interpretation, with source versions and access notes attached.

Download the evidence ↗