Nature Medicine: fully on-premise clinical agent scores 90.04% on a seven-disease MIMIC-IV benchmark beneficial
- The paper, published in Nature Medicine on 15 September 2026, reports that a fully on-premise clinical agent "achieved 90.04% accuracy on a seven-disease task and 83.8% accuracy on a four-disease task" across two MIMIC-IV-derived benchmarks. On the primary benchmark MIRA-v2, Qwen-3.5 reached 90.0%, GLM-5 89.7%, GLM-4.5-Air 88.4% and GPT-OSS 85.3%, against a cloud baseline of GPT-5.2 at 90.7% — "The best on-premise model was, therefore, within 0.7 percentage points of the cloud baseline."
- On the CDM benchmark (four abdominal categories, n = 2,400) Qwen-3.5 scored 83.8% and GLM-4.5-Air 81.2%; the paper states "The highest previously reported open-weight result on this benchmark was 70.5% (Gemma-3)."
- The authors report that behavioural consistency across repeated runs discriminated correct from incorrect diagnoses better than the model's own probability score (AUC = 0.860 versus 0.747), and that at a consistency threshold of 0.90, "49.4% of cases were retained at 98.9% diagnostic accuracy" — 272 cases routed to autonomous handling, with three errors among them.
- The paper states its own limits plainly: both primary benchmarks derive from MIMIC-IV and "a single-institution data ecology"; the evaluation is text-only; consistency thresholds "must be calibrated to the deployment configuration"; and "all evaluations were retrospective simulations", with prospective studies and bias audits still required.