Benchmark analysis / 5 min read

DDXPlus: a correct top diagnosis can hide a weak differential

Interpret differential recall, precision and pathology accuracy as different diagnostic targets.

The short answer

DDXPlus supplies a synthetic patient’s underlying pathology and a reference differential diagnosis. That creates two related but different targets: identify the simulated condition and preserve the appropriate alternatives. Our analysis uses the original baseline comparison to explain why a strong top-ranked prediction need not imply broad differential coverage. The benchmark remains a simulator-based research task, not evidence of performance on a prospective patient cohort.

Decide which target matters before selecting a metric

Top-pathology accuracy asks whether the first predicted condition matches the pathology used to synthesize the patient. Differential recall asks how much of the reference set was retained. Differential precision asks how much of the predicted set belongs to that reference. Those measurements can move in different directions because they reward different properties of the output.

An abstract example makes the distinction clear. Suppose the reference contains labels A, B and C, while a prediction gives A only. If A is the underlying pathology, the top prediction can be correct while two reference alternatives are missing. This is our illustrative arithmetic example, not a clinical case or reported experiment. It explains why the selected result panel retains the training target beside each method.

Apply the published threshold before comparing sets

The original scoring procedure removes differential entries whose probability is at or below 0.01. That operation turns a probability distribution into the set used for several metrics. A different threshold changes the set even if the model’s probabilities are unchanged. The threshold therefore belongs in the experiment configuration rather than a hidden plotting script.

Our recommended audit preserves both the original distribution and the post-threshold set. Review conditions just above and below the boundary, and distinguish a threshold change from a model change. If a team explores alternative thresholds, present the analysis as a sensitivity study with its own purpose. Do not silently substitute a threshold that improves the final reported F1 after inspecting the test results.

Keep information gathering in the picture

DDXPlus interactions begin with initial evidence and allow collection of structured symptoms and antecedents. The benchmark measures evidence collection and interaction length alongside diagnostic outputs. A system that reaches an answer quickly and a system that asks many questions have different interaction behavior even when their final pathology predictions agree.

Our interpretation is that the relevant tradeoff depends on the intended research question. More questions may recover additional evidence, but a raw turn count does not establish that the extra questions were useful. Inspect what was requested, what was learned and when the system stopped. The source supports controlled simulator comparisons; any inference about real patient burden would require additional evidence beyond these interaction metrics.

Respect the simulator’s scope

The original population is synthesized from a knowledge base and rule-based diagnostic system over selected pathologies. Large sample size can make a controlled experiment convenient, but repeating combinations generated by the same rules does not automatically add new disease mechanisms or repair missing rules. The synthetic origin should remain explicit in every description.

A useful report distinguishes benchmark agreement from clinical validity. Record language release, evidence schema, split, training target and scoring threshold. Then state which failure mode the experiment investigates, such as missing alternatives or inefficient evidence collection. This makes DDXPlus analytically useful without treating a simulator’s reference distribution as an independently observed clinical consensus. The study’s measured results remain credited to its authors.

References & further reading

These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.

  1. DDXPlus: A New Dataset For Automatic Medical Diagnosis ↗Fansi Tchango et al. / NeurIPS 2022. Synthetic population construction, evidence taxonomy and differential-diagnosis baselines.
  2. DDXPlus official repository ↗Mila / DDXPlus authors. Data schema and French/English release relationship.
  3. DDXPlus Dataset (English) ↗Fansi Tchango et al. / Figshare. Official English release, provenance and CC BY 4.0 license.

Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.

Continue reading.

All guides →