# Medical Evals > Independent analysis of DDXPlus and MedR-Bench: synthetic differentials, case-report reasoning, examination requests, oracle conditions, metric denominators and published results. DDXPlus separates a simulated pathology from its reference differential. MedR-Bench separates examination gathering, diagnosis, treatment and written reasoning. We analyze those task boundaries and the original scoring rules, with selected version-matched historical results. Our coverage explorer and guides are original analytical resources. The benchmarks and experiments belong to their credited authors; this site has not run the displayed models or created a new patient dataset. ## Provenance Independent analysis published by Arcophos. Benchmark creation belongs to the credited authors. Result rows are selected paper-reported measurements with their source versions and evaluation conditions, not new Arcophos runs or a live leaderboard. ## Benchmark dossiers - [DDXPlus](https://medicalevals.com/benchmarks/ddxplus/): The most likely disease and the useful differential are different targets. Source version: 2022 benchmark; English release tracked separately. - [MedR-Bench](https://medicalevals.com/benchmarks/medr-bench/): Supplying the examinations changes what a diagnosis score means. Source version: Nature Communications 2025, seven-model study. ## Original analyses - [DDXPlus: a correct top diagnosis can hide a weak differential](https://medicalevals.com/guides/ddxplus-differential-versus-top-one/): Interpret differential recall, precision and pathology accuracy as different diagnostic targets. - [MedR-Bench: read the gap between requested and supplied examinations](https://medicalevals.com/guides/medr-bench-oracle-gap/): Understand what changes between one-turn, free-turn and oracle diagnosis conditions. - [MedR-Bench reasoning scores: three denominators, three questions](https://medicalevals.com/guides/medr-reasoning-metric-denominators/): Analyze efficiency, factuality and completeness without treating written explanations as hidden model reasoning. ## Inspect the evidence - [Evidence JSON](https://medicalevals.com/evidence.json): Task definitions, dataset facts, scoring rules, source-version results, our interpretations, and reference IDs. - [Sources](https://medicalevals.com/sources/): Original papers and repositories with evidence locators. - [Editorial method](https://medicalevals.com/methodology/): Source reconciliation and interpretation boundaries. - [About](https://medicalevals.com/about/): Ownership and corrections. Analysis updated: 2026-09-28