Independent benchmark analysis / 2022 benchmark; English release tracked separately

DDXPlus

The most likely disease and the useful differential are different targets.

DDXPlus evaluates diagnostic information gathering against synthetic patients with both an underlying pathology and a reference differential. This enables a revealing comparison: a model can improve its top-ranked diagnosis while covering less of the differential. We interpret the original baseline results through that target distinction, preserve the probability threshold used to define a differential, and keep the synthetic population explicit. Our analysis focuses on evaluation design rather than proposing a clinical diagnostic tool. The benchmark’s scale and structured evidence make it useful for controlled experiments, but they do not create prospective evidence about real patients.

01 / What is being tested?

The task, before the score.

input
Initial evidence and subsequent responses from a synthetic patient simulator
output
Evidence requests and a pathology or differential distribution
unit
One synthetic patient interaction
setting
Original French release; rule-based simulator reference, with mean results over three runs

Data origin. Patients synthesized from a proprietary medical knowledge base and commercial rule-based diagnosis system. [1][2][3]

Synthetic population
Approximately 1.3 million

Paper-rounded total; do not describe these as real patient records.

Abstract; §3.4 [1]
Pathologies
49

Selected around cough, sore throat or breathing presentations.

§3 [1]
Evidence variables
223

110 symptoms and 113 antecedents in the original paper.

Table 1 [1]
Partition rule
80% / 10% / 10%

Training / validation / test, stratified on simulated pathology. Exact release row totals not independently counted here.

§3.4 [1]
Differential threshold
Probability > 0.01

Both predicted and reference differentials drop entries at or below 0.01.

Appendix H [1]
  1. 01

    Initialize a synthetic patient

    Start with the supplied initial evidence.

    [2]
  2. 02

    Collect structured evidence

    Ask about symptoms or antecedents.

    [2]
  3. 03

    Stop and predict

    Return an underlying pathology or differential distribution.

    [1]
  4. 04

    Apply the differential threshold

    Score coverage after filtering low-probability entries.

    [1]

Dataset anatomy

Evidence types in the original benchmark

Binary

Original paper definition.

208 evidence variables[1]
Categorical

Original paper definition.

10 evidence variables[1]
Multi-choice

Original paper definition.

5 evidence variables[1]

These are evidence variables, not patients or mutually exclusive findings within one case. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Differential recall, precision, F1 and pathology accuracy

Higher coverage/accuracy is better; interaction length is contextual

The benchmark distinguishes covering the reference differential from placing the simulated pathology first.

Scoring definition

DDR = mean(|prediction ∩ reference| / |reference|); DDP = mean(|prediction ∩ reference| / |prediction|)

Keep probability threshold, dialogue policy and training target fixed when comparing. [1]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Training target changes differential coverage

Original 2022 test; three-run means; original French data and Appendix H threshold.

Differential F1 · %
050100
Reported
AARLC, differential targetReported 95% CI half-width: 6.82 percentage points.
78.24%
AARLC, pathology targetReported 95% CI half-width: 0.38 points.
31.28%
BASD, differential targetReported 95% CI half-width: 1.57 points.
83.69%
BASD, pathology targetReported 95% CI half-width: 0.29 points.
31.31%

Historical paper-reported means. They are simulator agreement metrics, not clinical diagnostic accuracy.

Source: Table 3 [1]

04 / Our original analysis

What follows from the design?

01

Optimize the target you intend to measure

Published evidence

The paper separately evaluates simulated pathology and differential coverage. [1]

Our interpretation

Our inference: a single top-1 number can hide whether plausible alternatives survive evidence gathering.

02

Thresholding changes the set being judged

Published evidence

Differential entries at or below 0.01 are removed. [1]

Our interpretation

Our inference: the threshold belongs in a result identifier; changing it can move precision and recall without changing the underlying probabilities.

03

A large simulator can have a narrow support

Published evidence

The synthetic patients come from selected pathology and evidence rules. [1][3]

Our interpretation

Our inference: more sampled patients need not add new mechanisms or eliminate systematic omissions in those rules.

05 / Scope of the evidence

Where this benchmark stops.

Synthetic reference

Agreement with a simulator does not establish agreement with a real-world clinical differential. [3]

Selected pathology universe

The benchmark is not a comprehensive disease catalog. [1]

Version and language

The authors state English and French versions represent the same data, but release identifiers still matter. [2]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Public release
License
CC BY 4.0
Conditions
Published data are synthetic; cite the precise language and Figshare version.
[3]

Evidence trail

Read the originals.

  1. DDXPlus: A New Dataset For Automatic Medical Diagnosis ↗

    Fansi Tchango et al. / NeurIPS 2022. Synthetic population construction, evidence taxonomy and differential-diagnosis baselines.

  2. DDXPlus official repository ↗

    Mila / DDXPlus authors. Data schema and French/English release relationship.

  3. DDXPlus Dataset (English) ↗

    Fansi Tchango et al. / Figshare. Official English release, provenance and CC BY 4.0 license.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗