{"publication":"Medical Evals","url":"https://medicalevals.com","publisher":"Arcophos","updated":"2026-09-28","provenance":"Independent analytical publication. Benchmark creation and experimental results belong to their cited authors. Reported results are source-version snapshots, not new Arcophos runs or a live leaderboard.","benchmarks":[{"slug":"ddxplus","name":"DDXPlus","shortName":"DDXPlus","version":"2022 benchmark; English release tracked separately","creators":"Arsene Fansi Tchango, Rishab Goel, Zhi Wen, Julien Martel and Joumana Ghosn","paperDate":"2022","headline":"The most likely disease and the useful differential are different targets.","summary":"DDXPlus evaluates diagnostic information gathering against synthetic patients with both an underlying pathology and a reference differential. This enables a revealing comparison: a model can improve its top-ranked diagnosis while covering less of the differential. We interpret the original baseline results through that target distinction, preserve the probability threshold used to define a differential, and keep the synthetic population explicit. Our analysis focuses on evaluation design rather than proposing a clinical diagnostic tool. The benchmark’s scale and structured evidence make it useful for controlled experiments, but they do not create prospective evidence about real patients.","task":{"input":"Initial evidence and subsequent responses from a synthetic patient simulator","output":"Evidence requests and a pathology or differential distribution","unit":"One synthetic patient interaction","setting":"Original French release; rule-based simulator reference, with mean results over three runs"},"dataOrigin":"Patients synthesized from a proprietary medical knowledge base and commercial rule-based diagnosis system.","facts":[{"label":"Synthetic population","value":"Approximately 1.3 million","detail":"Paper-rounded total; do not describe these as real patient records.","sourceIds":["ddx-paper"],"locator":"Abstract; §3.4"},{"label":"Pathologies","value":"49","detail":"Selected around cough, sore throat or breathing presentations.","sourceIds":["ddx-paper"],"locator":"§3"},{"label":"Evidence variables","value":"223","detail":"110 symptoms and 113 antecedents in the original paper.","sourceIds":["ddx-paper"],"locator":"Table 1"},{"label":"Partition rule","value":"80% / 10% / 10%","detail":"Training / validation / test, stratified on simulated pathology. Exact release row totals not independently counted here.","sourceIds":["ddx-paper"],"locator":"§3.4"},{"label":"Differential threshold","value":"Probability > 0.01","detail":"Both predicted and reference differentials drop entries at or below 0.01.","sourceIds":["ddx-paper"],"locator":"Appendix H"}],"metric":{"name":"Differential recall, precision, F1 and pathology accuracy","description":"The benchmark distinguishes covering the reference differential from placing the simulated pathology first.","formula":"DDR = mean(|prediction ∩ reference| / |reference|); DDP = mean(|prediction ∩ reference| / |prediction|)","direction":"Higher coverage/accuracy is better; interaction length is contextual","comparability":"Keep probability threshold, dialogue policy and training target fixed when comparing.","sourceIds":["ddx-paper"]},"workflow":[{"label":"Initialize a synthetic patient","detail":"Start with the supplied initial evidence.","sourceIds":["ddx-repo"]},{"label":"Collect structured evidence","detail":"Ask about symptoms or antecedents.","sourceIds":["ddx-repo"]},{"label":"Stop and predict","detail":"Return an underlying pathology or differential distribution.","sourceIds":["ddx-paper"]},{"label":"Apply the differential threshold","detail":"Score coverage after filtering low-probability entries.","sourceIds":["ddx-paper"]}],"slices":[{"label":"Binary","value":208,"unit":"evidence variables","detail":"Original paper definition.","sourceIds":["ddx-paper"]},{"label":"Categorical","value":10,"unit":"evidence variables","detail":"Original paper definition.","sourceIds":["ddx-paper"]},{"label":"Multi-choice","value":5,"unit":"evidence variables","detail":"Original paper definition.","sourceIds":["ddx-paper"]}],"sliceTitle":"Evidence types in the original benchmark","sliceNote":"These are evidence variables, not patients or mutually exclusive findings within one case.","results":[{"id":"differential-f1","title":"Training target changes differential coverage","metric":"Differential F1","unit":"%","lower":0,"upper":100,"scope":"Original 2022 test; three-run means; original French data and Appendix H threshold.","sourceIds":["ddx-paper"],"locator":"Table 3","rows":[{"label":"AARLC, differential target","value":78.24,"display":"78.24%","detail":"Reported 95% CI half-width: 6.82 percentage points."},{"label":"AARLC, pathology target","value":31.28,"display":"31.28%","detail":"Reported 95% CI half-width: 0.38 points."},{"label":"BASD, differential target","value":83.69,"display":"83.69%","detail":"Reported 95% CI half-width: 1.57 points."},{"label":"BASD, pathology target","value":31.31,"display":"31.31%","detail":"Reported 95% CI half-width: 0.29 points."}],"note":"Historical paper-reported means. They are simulator agreement metrics, not clinical diagnostic accuracy."}],"analysis":[{"heading":"Optimize the target you intend to measure","evidence":"The paper separately evaluates simulated pathology and differential coverage.","interpretation":"Our inference: a single top-1 number can hide whether plausible alternatives survive evidence gathering.","sourceIds":["ddx-paper"]},{"heading":"Thresholding changes the set being judged","evidence":"Differential entries at or below 0.01 are removed.","interpretation":"Our inference: the threshold belongs in a result identifier; changing it can move precision and recall without changing the underlying probabilities.","sourceIds":["ddx-paper"]},{"heading":"A large simulator can have a narrow support","evidence":"The synthetic patients come from selected pathology and evidence rules.","interpretation":"Our inference: more sampled patients need not add new mechanisms or eliminate systematic omissions in those rules.","sourceIds":["ddx-paper","ddx-data"]}],"limitations":[{"title":"Synthetic reference","detail":"Agreement with a simulator does not establish agreement with a real-world clinical differential.","sourceIds":["ddx-data"]},{"title":"Selected pathology universe","detail":"The benchmark is not a comprehensive disease catalog.","sourceIds":["ddx-paper"]},{"title":"Version and language","detail":"The authors state English and French versions represent the same data, but release identifiers still matter.","sourceIds":["ddx-repo"]}],"access":{"status":"Public release","license":"CC BY 4.0","restrictions":"Published data are synthetic; cite the precise language and Figshare version.","url":"https://figshare.com/articles/dataset/DDXPlus_Dataset_English_/22687585","sourceIds":["ddx-data"]},"sourceIds":["ddx-paper","ddx-repo","ddx-data"]},{"slug":"medr-bench","name":"MedR-Bench","shortName":"MedR-Bench","version":"Nature Communications 2025, seven-model study","creators":"Pengcheng Qiu, Chaoyi Wu and collaborators / Shanghai Jiao Tong University and Shanghai AI Laboratory","paperDate":"2025-11-06","headline":"Supplying the examinations changes what a diagnosis score means.","summary":"MedR-Bench reorganizes published case reports into staged clinical tasks. It evaluates examination recommendations, diagnoses and treatment plans, while a separate automated evaluator scores written reasoning. Our analysis follows the information made available at each stage: one-turn diagnosis and oracle diagnosis are deliberately different conditions. We reproduce selected version-matched measurements and explain why factual reasoning, complete reasoning and a correct final answer must remain distinct. The source cases are real published reports, but the interactions and scoring are reconstructed research procedures rather than a prospective clinical study. This site contributes interpretation, not new experimental results.","task":{"input":"Structured case information, with examinations withheld or supplied according to setting","output":"Requested examinations, diagnosis or treatment plan plus written rationale","unit":"One case-report-derived patient case","setting":"One-turn, free-turn and oracle conditions; reasoning judged through an agentic GPT-4o-based pipeline"},"dataOrigin":"PMC Open Access case reports reorganized into structured cases using GPT-4o.","facts":[{"label":"Structured cases","value":"1,453","detail":"957 diagnosis cases plus 496 treatment cases.","sourceIds":["medr-paper"],"locator":"Patient cases"},{"label":"Rare-disease cases","value":"656","detail":"491 diagnosis and 165 treatment cases; these are subsets, not extra cases.","sourceIds":["medr-paper"],"locator":"Patient cases"},{"label":"Body systems","value":"13","detail":"Source taxonomy coverage.","sourceIds":["medr-paper"],"locator":"Abstract"},{"label":"Clinical stages","value":"3","detail":"Examination recommendation, diagnosis and treatment.","sourceIds":["medr-paper"],"locator":"Evaluation settings"},{"label":"Study version","value":"7 models","detail":"The journal study expands the five-model March 2025 preprint. Use version-matched results.","sourceIds":["medr-paper"],"locator":"Abstract; evaluated models"}],"metric":{"name":"Outcome accuracy and written-reasoning dimensions","description":"Outcome matching uses GPT-4o; efficiency, factuality and completeness use distinct step denominators. Completeness is unavailable for examination recommendation because the case reports rarely document why examinations were selected.","formula":"Efficiency = effective / generated steps; factuality = correct effective / effective steps; completeness = covered reference / reference steps","direction":"Higher is better within the same condition","comparability":"Written rationale evaluation is not direct observation of a model’s internal causal reasoning. Freeze judge, retrieval and prompts.","sourceIds":["medr-paper","medr-repo"]},"workflow":[{"label":"Choose information availability","detail":"Use one-turn, free-turn or oracle access to examination evidence.","sourceIds":["medr-paper"]},{"label":"Generate clinical outputs","detail":"Request examinations or produce a diagnosis/treatment plan.","sourceIds":["medr-repo"]},{"label":"Evaluate outcomes","detail":"Match outputs to the versioned references with the source judge.","sourceIds":["medr-paper"]},{"label":"Evaluate written reasoning","detail":"Decompose steps, verify factuality and match reference steps.","sourceIds":["medr-paper"]}],"slices":[{"label":"Diagnosis","value":957,"unit":"cases","detail":"Contains 491 rare-disease cases.","sourceIds":["medr-paper"]},{"label":"Treatment","value":496,"unit":"cases","detail":"Contains 165 rare-disease cases.","sourceIds":["medr-paper"]}],"sliceTitle":"Case types, without double-counting rare diseases","sliceNote":"Rare-disease labels overlap these task subsets. They must not be added as a third disjoint group.","results":[{"id":"one-turn","title":"Diagnosis after one examination turn","metric":"Accuracy","unit":"%","lower":0,"upper":100,"scope":"2025 journal supplement, all 957 diagnosis cases; one-turn examination condition.","sourceIds":["medr-supp"],"locator":"Supplementary Table 3","rows":[{"label":"OpenAI-o3-mini","value":64.99,"display":"64.99%","detail":"95% CI 61.97–68.02; o3-mini-2025-01-31."},{"label":"Gemini-2.0-FT","value":68.55,"display":"68.55%","detail":"95% CI 65.60–71.49; flash-thinking-exp-01-21."},{"label":"DeepSeek-R1","value":71.79,"display":"71.79%","detail":"95% CI 68.93–74.64; 671B model."}],"note":"Paper-reported values, not our runs. Source uses two-sided 95% intervals. Generation and judge configurations are documented in Supplement §A.2."},{"id":"oracle","title":"Diagnosis with all reference examinations","metric":"Accuracy","unit":"%","lower":0,"upper":100,"scope":"Same journal study and diagnosis cases; oracle evidence condition.","sourceIds":["medr-supp"],"locator":"Supplementary Table 3","rows":[{"label":"OpenAI-o3-mini","value":83.91,"display":"83.91%","detail":"95% CI 81.58–86.24."},{"label":"Gemini-2.0-FT","value":86.83,"display":"86.83%","detail":"95% CI 84.69–88.98."},{"label":"DeepSeek-R1","value":89.76,"display":"89.76%","detail":"95% CI 87.84–91.68."}],"note":"The evidence condition changes. This panel is not a matched deployment comparison or a new-model release trend."}],"analysis":[{"heading":"Information gathering is separately measurable","evidence":"The same benchmark includes one-turn and oracle diagnosis conditions.","interpretation":"Our inference: report both before treating a high oracle score as evidence that a system knows which examinations to request.","sourceIds":["medr-paper","medr-supp"]},{"heading":"Reasoning dimensions have different denominators","evidence":"Factuality considers effective steps; completeness considers reference steps.","interpretation":"Our inference: a short entirely factual rationale can still omit most required reasoning, so the metrics should not be collapsed.","sourceIds":["medr-paper"]},{"heading":"Case reports create a publication filter","evidence":"Structured cases originate in published case reports.","interpretation":"Our inference: their disease mix and narrative completeness need not represent consecutive patients in a clinic.","sourceIds":["medr-paper"]},{"heading":"A judge is part of the measurement","evidence":"The evaluator uses language models and external search.","interpretation":"Our inference: preserve retrieval snapshots and judge configuration when a result needs to be audited later.","sourceIds":["medr-paper","medr-repo"]}],"limitations":[{"title":"Reconstructed patient interactions","detail":"The benchmark uses staged case information and an agentic patient role, not a prospective clinical encounter.","sourceIds":["medr-paper"]},{"title":"Written reasoning reference","detail":"A published discussion cannot document every valid reasoning route or expose hidden model computation.","sourceIds":["medr-paper"]},{"title":"Automated judgment dependency","detail":"Semantic matching, web retrieval and judge behavior can affect the scores.","sourceIds":["medr-paper"]}],"access":{"status":"Public repository and supplements","license":"Repository identifies CC BY-SA; version not specified in its short LICENSE","restrictions":"Consult the creators for exact reuse terms and original article rights before redistributing cases.","url":"https://github.com/MAGIC-AI4Med/MedRBench","sourceIds":["medr-repo"]},"sourceIds":["medr-paper","medr-supp","medr-repo"]}],"explorer":{"kind":"coverage","title":"Map the diagnostic target and information path","intro":"Compare actual simulator and case-report tasks. Filter by diagnostic target, information gathering, clinical stage or written-reasoning evaluation.","caution":"This is our independent analytical map of published tasks. It does not execute an evaluation, predict a model’s performance or establish clinical benefit.","sourceIds":["ddx-paper","ddx-repo","ddx-data","medr-paper","medr-supp","medr-repo"],"rows":[{"label":"DDXPlus: top pathology","category":"Diagnostic target","input":"Collected simulator evidence","output":"Most likely pathology","metric":"GTPA@1","constraint":"Synthetic underlying pathology is the target.","benchmarkSlug":"ddxplus","sourceIds":["ddx-paper"]},{"label":"DDXPlus: differential coverage","category":"Diagnostic target","input":"Collected simulator evidence","output":"Set of plausible pathologies","metric":"DDR / DDP / DDF1","constraint":"Apply the published probability threshold before set comparison.","benchmarkSlug":"ddxplus","sourceIds":["ddx-paper"]},{"label":"DDXPlus: evidence gathering","category":"Information gathering","input":"Initial symptom and interactive evidence","output":"Evidence requests","metric":"Positive evidence recall; interaction length","constraint":"A longer interaction is not automatically better.","benchmarkSlug":"ddxplus","sourceIds":["ddx-paper"]},{"label":"MedR: one-turn examinations","category":"Information gathering","input":"Case without ancillary test findings","output":"Recommended examination list","metric":"Precision and recall","constraint":"References describe examinations present in the source case report.","benchmarkSlug":"medr-bench","sourceIds":["medr-paper"]},{"label":"MedR: free-turn examinations","category":"Information gathering","input":"Case with iterative examination access","output":"Sequence of requests","metric":"Precision and recall","constraint":"The model chooses when to stop; patient responses are agent-mediated.","benchmarkSlug":"medr-bench","sourceIds":["medr-paper"]},{"label":"MedR: one-turn diagnosis","category":"Clinical stage","input":"Case plus once-requested examinations","output":"Diagnosis","metric":"Semantic accuracy","constraint":"Information availability depends on the preceding request.","benchmarkSlug":"medr-bench","sourceIds":["medr-paper"]},{"label":"MedR: oracle diagnosis","category":"Clinical stage","input":"All reference examinations","output":"Diagnosis","metric":"Semantic accuracy","constraint":"Does not test selecting the examinations.","benchmarkSlug":"medr-bench","sourceIds":["medr-paper"]},{"label":"MedR: oracle treatment","category":"Clinical stage","input":"Complete case including diagnosis","output":"Treatment plan","metric":"Semantic plan accuracy","constraint":"Case-report treatment is the reference; not a prospective intervention comparison.","benchmarkSlug":"medr-bench","sourceIds":["medr-paper"]},{"label":"MedR: written-reasoning evaluation","category":"Reasoning","input":"Generated rationale and reference reasoning","output":"Judged steps","metric":"Efficiency, factuality, completeness","constraint":"Measures expressed rationale with an automated evaluator. Examination recommendation lacks reference reasoning for completeness scoring.","benchmarkSlug":"medr-bench","sourceIds":["medr-paper","medr-repo"]}],"parameters":[]},"references":[{"id":"ddx-paper","title":"DDXPlus: A New Dataset For Automatic Medical Diagnosis","organization":"Fansi Tchango et al. / NeurIPS 2022","url":"https://arxiv.org/pdf/2205.09148","note":"Synthetic population construction, evidence taxonomy and differential-diagnosis baselines.","locator":"§3–4; Tables 1–3; Appendix H","version":"2022"},{"id":"ddx-repo","title":"DDXPlus official repository","organization":"Mila / DDXPlus authors","url":"https://github.com/mila-iqia/ddxplus","note":"Data schema and French/English release relationship.","locator":"Availability; dataset documentation","version":"Accessed 2026-09-28"},{"id":"ddx-data","title":"DDXPlus Dataset (English)","organization":"Fansi Tchango et al. / Figshare","url":"https://figshare.com/articles/dataset/DDXPlus_Dataset_English_/22687585","note":"Official English release, provenance and CC BY 4.0 license.","locator":"Version 2; Licence","version":"2026-03-09"},{"id":"medr-paper","title":"Quantifying the reasoning abilities of LLMs on clinical cases","organization":"Qiu et al. / Nature Communications","url":"https://www.nature.com/articles/s41467-025-64769-1","note":"Peer-reviewed seven-model study, case counts, written-reasoning metrics and judge implementation.","locator":"Patient cases; Evaluation metrics; Methods","version":"2025-11-06"},{"id":"medr-supp","title":"MedR-Bench Supplementary Information","organization":"Qiu et al. / Nature Communications","url":"https://media.springernature.com/original/springer-static/esm/art%3A10.1038%2Fs41467-025-64769-1/MediaObjects/41467_2025_64769_MOESM1_ESM.pdf","note":"Version-matched results with confidence intervals and generation parameters.","locator":"Supplementary Table 3; §A.2","version":"2025-11-06"},{"id":"medr-repo","title":"MedR-Bench official code and data","organization":"MAGIC-AI4Med","url":"https://github.com/MAGIC-AI4Med/MedRBench","note":"Source code, dataset and saved model responses; repository license identifies CC BY-SA without a version.","locator":"README; LICENSE","version":"Accessed 2026-09-28"}]}