Benchmark dossiers / independent analysis

The benchmark behind the number.

Inspect what each benchmark asks, how its score is produced, and which conclusions its design can support. These are original analyses by Arcophos of the benchmark authors’ work. Any reproduced measurements are labeled with their published source and version.

Dossier01

2022 benchmark; English release tracked separately

DDXPlus ↗

The most likely disease and the useful differential are different targets.

UnitOne synthetic patient interactionMeasureDifferential recall, precision, F1 and pathology accuracy
Dossier02

Nature Communications 2025, seven-model study

MedR-Bench ↗

Supplying the examinations changes what a diagnosis score means.

UnitOne case-report-derived patient caseMeasureOutcome accuracy and written-reasoning dimensions
What we contribute. Source reconciliation, task and metric interpretation, and tools that make assumptions inspectable. We do not claim to have created these benchmarks or run the reported models.