Topics / topic

Healthcare

10 items across 4 editions · appeared in the last 4 editions in a row. First seen Sat 12 Sep, last seen Tue 15 Sep. Traced across 1 weekly review.

How this story has evolved

From the week in review: the connections, developments and open questions filed under Healthcare, newest week first.

Week of 7–13 September 2026

Development · Tue 8 Sep
DeepMind releases AlphaGenome Atlas: precomputed predictions for 9 billion single-letter DNA changes in a 1-petabyte dataset

Google DeepMind published AlphaGenome Atlas on 8 September, with predictions for the effects of "9 billion single-nucleotide variants — every single-letter change possible" in the human genome, held in "a massive 1-petabyte dataset, more than 30 times larger than the AlphaFold Database".

Tuesday, 15 September 2026

Nature Medicine: fully on-premise clinical agent scores 90.04% on a seven-disease MIMIC-IV benchmark beneficial

  • The paper, published in Nature Medicine on 15 September 2026, reports that a fully on-premise clinical agent "achieved 90.04% accuracy on a seven-disease task and 83.8% accuracy on a four-disease task" across two MIMIC-IV-derived benchmarks. On the primary benchmark MIRA-v2, Qwen-3.5 reached 90.0%, GLM-5 89.7%, GLM-4.5-Air 88.4% and GPT-OSS 85.3%, against a cloud baseline of GPT-5.2 at 90.7% — "The best on-premise model was, therefore, within 0.7 percentage points of the cloud baseline."
  • On the CDM benchmark (four abdominal categories, n = 2,400) Qwen-3.5 scored 83.8% and GLM-4.5-Air 81.2%; the paper states "The highest previously reported open-weight result on this benchmark was 70.5% (Gemma-3)."
  • The authors report that behavioural consistency across repeated runs discriminated correct from incorrect diagnoses better than the model's own probability score (AUC = 0.860 versus 0.747), and that at a consistency threshold of 0.90, "49.4% of cases were retained at 98.9% diagnostic accuracy" — 272 cases routed to autonomous handling, with three errors among them.
  • The paper states its own limits plainly: both primary benchmarks derive from MIMIC-IV and "a single-institution data ecology"; the evaluation is text-only; consistency thresholds "must be calibrated to the deployment configuration"; and "all evaluations were retrospective simulations", with prospective studies and bias audits still required.

Audit of 26 language models finds 55.4% of generated biomedical references fabricated harmfulPreprint

  • arXiv:2609.14988, "Biomedical Reference Generation Remains Unreliable across 26 Large Language Models", submitted 14 September 2026 by Maxim Topaz and colleagues, prompted "26 language models from eight developers (2023 to 2026) to supply a missing reference for each of 69 biomedical passages across ten domains".
  • The paper reports: "Across all models, 55.4% of responses were fabricated and 14.9% were correct in every field." Fabrication "ranged from 10.2% (Claude Opus 4.8, which declined 52.1% of prompts) to 98.4% (Ministral 3B, which produced no verifiable reference)".
  • Among models first released in 2026, the paper reports fabricated and all-fields-correct proportions of "35.3% and 31.8%, respectively". GPT-5.5 "was correct in every field in 48.1%", and Claude Opus 4.6 and Claude Sonnet 4.5 produced similar proportions of verifiable references (77.6% and 76.6%) but were correct in every evaluated field in 54.6% and 19.9% of responses.
  • The authors conclude that "no model was correct in every evaluated bibliographic field in more than 54.6% of responses" and that "references produced with model assistance require verification before use". The paper is a preprint and has not been peer reviewed.

Documents from EFF FOIA suit show Medicare's AI prior-authorisation pilot launched on untested software harmfulUpdate

  • STAT reported on 15 September that the rollout of the Wasteful and Inappropriate Service Reduction model, or WISeR, "was hasty and error-ridden, according to more than a thousand pages of recently released documents and data" obtained by the Electronic Frontier Foundation through a Freedom of Information Act lawsuit against the Centers for Medicare and Medicaid Services.
  • STAT reports the pilot launched in January, requires prior approval for certain procedures and products including skin substitutes and epidural injections for pain management, operates in New Jersey, Ohio, Oklahoma, Texas, Arizona and Washington, and will run until 2031. STAT reports the documents show one WISeR vendor warned CMS that it was unrealistic to expect a working product by the launch date the agency wanted.
  • The EFF's own analysis of the same records, published 8 September, reports that one prior authorisation request went unanswered for 83 days against a 72-hour standard, that two vendors alone denied over 20,000 requests in the first three months, that Virtix denied more requests than it approved in that period, and that low quality scores reduce vendor payments by only 5–10%. EFF quotes Innovaccer telling CMS about a month before launch that "auto-affirming is the only path available".
  • The remainder of the STAT article is paywalled, so the figures in the previous bullet come from the EFF analysis rather than from STAT. CMS has not published a response to the released records.

Monday, 14 September 2026

Nature Medicine: AI support raised physicians' lung-cancer disease-control sensitivity from 0.72 to 0.87 in a 2,396-patient study beneficial

  • "Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC", published in Nature Medicine on 13 September, reports on I³LUNG (NCT05537922), which enrolled 2,396 patients with stage IIIC–IVB non-small cell lung cancer treated with immunotherapy between September 2012 and October 2023 across six centres in Italy, Greece, Germany, Spain, the USA and Israel.
  • In the usability study, 20 physicians — 10 lung expert oncologists and 10 non-experts — each assessed 10 real-world cases, totalling 200 assessments. With the explainable-AI tool, sensitivity for disease control rate rose from 0.72 (95% CI 0.64–0.80) to 0.87 (95% CI 0.79–0.92), P = 0.0011, and accuracy from 0.57 to 0.65, P = 0.0431, "at the expense of slightly lower specificity". Agreement between experts and non-experts rose from κ = 0.11 to κ = 0.48.
  • Clinical-and-blood-only models "achieved consistent performance across outcomes with area under the curve (AUC) up to 0.77 in the test (TEST) set" and "significantly surpassed PD-L1" and other standard scores in that set. The authors say a prospective validation in more than 2,000 patients is under way.
  • The paper is explicit about limits: performance fell in external validation to an "AUC range: 0.55–0.72", and the benefit of adding CT, pathology and genomics "remains uncertain, not translated in TEST and EXVAL". Physicians adopted correct AI suggestions 74.5% of the time, but experts followed incorrect ones more often than non-experts, 72.2% versus 63.6%.

Harvard-led team releases a fine-tuned physician-level judge and a 9,217-score benchmark for grading medical AI beneficialPreprint

  • "Scaling Clinical Judgment to Evaluate Medical AI" (arXiv 2609.12822, submitted 11 September, announced 14 September) has 17 authors including Thomas A. Buckley, Adam Rodman and Arjun K. Manrai. It introduces PrecepTron, "an LLM fine-tuned for physician-level evaluation of open-ended responses", "trained using low-rank adaptation (LoRA) of a 32-billion-parameter model on a small number of physician examples".
  • The authors release GRAND-ROUNDS, "a new large-scale physician-annotated benchmark of 9,217 scores by 11 physicians across seven studies", and say all code, data and labels are freely available.
  • They report that frontier models used in typical "LLM-as-a-judge" setups "often disagree with physicians and with each other", and say they used PrecepTron to reproduce headline findings from five studies of LLMs for clinical care published in JAMA, Science and Nature Medicine "without new human grading".
  • Preprint, not peer reviewed. Reproducing published findings is not the same as validating clinical safety, and the abstract gives no figure for PrecepTron's agreement with physicians outside the seven studies it was built from.

Oxford study: rubric scoring leaves clinically relevant medical hallucinations undetected, often leaving scores unchanged mixedPreprint

  • "When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation" (arXiv 2609.12718, submitted 11 September, announced 14 September) has seven authors, all listed with University of Oxford affiliations in the HTML version. They build a taxonomy of medical hallucination types and a clinician-validated error-injection pipeline producing matched correct and error-injected responses.
  • The abstract states: "Across HealthBench, HealthBench Professional, and LiveMedBench, our clinically relevant hallucinations are missed by rubrics, often leaving scores unchanged." Rubrics "are most effective when explicitly checking facts, and are less effective for additional or unexpected errors they do not anticipate".
  • The authors report that a preliminary retrieval-based factuality check "recovers some of the rubric-blind errors", and conclude that "rubric scores alone are insufficient to establish clinical reliability".
  • Preprint, not peer reviewed. The hallucinations are injected by the authors rather than produced by a model in clinical use, and the abstract gives no figure for the share of injected errors that rubrics miss.

Sunday, 13 September 2026

UPenn preprint: self-supervised plasma proteomic model predicts 144 diseases across differing protein panels beneficialPreprint

  • The preprint, posted 12 September by Yonghyun Nam, Dokyoon Kim and colleagues at the University of Pennsylvania, reports a self-supervised model built on "53,014 participants in the UK Biobank Pharma Proteomics Project", covering 2,920-protein profiles and a predefined 1,460-protein subset, evaluated across 144 diseases.
  • Reported performance: "median AUC was 0.679 with comprehensive coverage and 0.637 when applied to partial-coverage representations"; retraining only the disease-specific models raised the partial-coverage median AUC to 0.673.
  • The authors report their protein-token risk scores exceeded coefficient-truncated LASSO by a median paired AUC difference of 0.027, and were comparable to LASSO refitted with outcome labels, a median difference of 0.003.
  • This is a preprint and has not been peer reviewed. A median AUC of 0.679 across 144 diseases is a population-level discrimination figure, not a clinical test, and the work is validated inside one cohort.

Preprint reports 99.0% cross-validated sensitivity separating early-stage ovarian cancer from controls in two small cohorts beneficialPreprintSingle source

  • Posted 12 September by Hongyi Zhou, Jean-Luc Chaubard, Benedict Benigno and Jeffrey Skolnick of Georgia Institute of Technology, OmicsIQ LLC and the Ovarian Cancer Institute, the preprint applies boosted decision tree classifiers to blood metabolomic data.
  • Cohorts are "91 serum samples (59 ovarian cancer, 32 healthy controls)" and "83 plasma samples (63 ovarian cancer, 20 healthy controls)". Reported mean cross-validated sensitivity and specificity are 99.0% and 99.8% in serum and 97.7% and 99.7% in plasma.
  • The authors report "239 concordantly altered annotated features spanning lipid, amino-acid, steroid, central-carbon, and redox metabolism", and propose the term "metabolomic Systemotype".
  • The results come from five-fold cross-validation repeated over 50 randomised rounds, not external validation, on fewer than 200 samples in total. Accuracy figures this high on cohorts this small do not establish screening performance in a general population, and the preprint has not been peer reviewed.

Saturday, 12 September 2026

Mayo Clinic study: AI reading of routine slides links tumour spatial pattern to 71% higher pancreatic cancer recurrence risk beneficial

  • The study, published in Clinical Cancer Research on 11 September as "Spatial Configuration of Pancreatic Cancer Is Associated with Disease Recurrence after Neoadjuvant Therapy and Curative-Intent Resection" (DOI 10.1158/1078-0432.ccr-25-4968), analysed tissue from 203 patients with pancreatic ductal adenocarcinoma.
  • The AI measured "tissue shape, fragmentation and the degree to which the cancer and stroma were intermixed" on standard pathology slides. "High-risk patients had a 71% higher adjusted risk of recurrence" in one model and "more than twice the adjusted risk" in another, while "the amount of residual cancer alone did not reliably separate patients at higher and lower risk".
  • High-risk spatial patterns "contained fewer immune cells within the cancer itself, with immune cells tending to collect around the tumor instead of entering it" — a mechanism, not just a correlation, and one that runs on slides hospitals already produce.
  • Dr Ryan Carr said "the results are promising but need to be confirmed in prospective studies before this approach could be used to inform clinical decision-making". Mayo Clinic's own newsroom page blocked our fetcher, so the figures above are as Medical Xpress reports them.

FDA clears Omniscient's Quicktome Deep Brain for mapping brain networks in deep brain stimulation planning Company claim

  • Omniscient announced on 11 September at 14:42 ET that Quicktome Deep Brain (DB) has received 510(k) clearance from the FDA, its third clearance.
  • The company says the workflow localises leads, segments deep brain structures and maps white matter tracts within brain networks, for deep brain stimulation used in Parkinson's disease, essential tremor, dystonia and other movement disorders.
  • The release states roughly 10,000 new DBS procedures are performed in the US each year — the scale against which any benefit would be measured.
  • CEO Stephen Scheeler is quoted: "This is our third FDA clearance, and each one advances the same strategy: build the connectomic platform that becomes the standard across neuroscience care." A 510(k) clearance establishes substantial equivalence to a predicate device, not clinical benefit; the release reports no outcome data.