Optimize the target you intend to measure
The paper separately evaluates simulated pathology and differential coverage. [1]
Our inference: a single top-1 number can hide whether plausible alternatives survive evidence gathering.
Independent benchmark analysis / 2022 benchmark; English release tracked separately
DDXPlus evaluates diagnostic information gathering against synthetic patients with both an underlying pathology and a reference differential. This enables a revealing comparison: a model can improve its top-ranked diagnosis while covering less of the differential. We interpret the original baseline results through that target distinction, preserve the probability threshold used to define a differential, and keep the synthetic population explicit. Our analysis focuses on evaluation design rather than proposing a clinical diagnostic tool. The benchmark’s scale and structured evidence make it useful for controlled experiments, but they do not create prospective evidence about real patients.
01 / What is being tested?
Data origin. Patients synthesized from a proprietary medical knowledge base and commercial rule-based diagnosis system. [1][2][3]
Paper-rounded total; do not describe these as real patient records.
Abstract; §3.4 [1]Selected around cough, sore throat or breathing presentations.
§3 [1]110 symptoms and 113 antecedents in the original paper.
Table 1 [1]Training / validation / test, stratified on simulated pathology. Exact release row totals not independently counted here.
§3.4 [1]Both predicted and reference differentials drop entries at or below 0.01.
Appendix H [1]Start with the supplied initial evidence.
[2]Ask about symptoms or antecedents.
[2]Return an underlying pathology or differential distribution.
[1]Score coverage after filtering low-probability entries.
[1]Dataset anatomy
Original paper definition.
Original paper definition.
Original paper definition.
These are evidence variables, not patients or mutually exclusive findings within one case. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Higher coverage/accuracy is better; interaction length is contextual
The benchmark distinguishes covering the reference differential from placing the simulated pathology first.
DDR = mean(|prediction ∩ reference| / |reference|); DDP = mean(|prediction ∩ reference| / |prediction|)
Keep probability threshold, dialogue policy and training target fixed when comparing. [1]
03 / Measured evidence
Paper-reported results / selected rows
Original 2022 test; three-run means; original French data and Appendix H threshold.
Historical paper-reported means. They are simulator agreement metrics, not clinical diagnostic accuracy.
Source: Table 3 [1]
04 / Our original analysis
The paper separately evaluates simulated pathology and differential coverage. [1]
Our inference: a single top-1 number can hide whether plausible alternatives survive evidence gathering.
Differential entries at or below 0.01 are removed. [1]
Our inference: the threshold belongs in a result identifier; changing it can move precision and recall without changing the underlying probabilities.
05 / Scope of the evidence
Agreement with a simulator does not establish agreement with a real-world clinical differential. [3]
The benchmark is not a comprehensive disease catalog. [1]
The authors state English and French versions represent the same data, but release identifiers still matter. [2]
Evidence trail
Fansi Tchango et al. / NeurIPS 2022. Synthetic population construction, evidence taxonomy and differential-diagnosis baselines.
Mila / DDXPlus authors. Data schema and French/English release relationship.
Fansi Tchango et al. / Figshare. Official English release, provenance and CC BY 4.0 license.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.