Topics / topic

AI For Science

8 items across 3 editions. First seen Sat 12 Sep, last seen Mon 14 Sep. Traced across 1 weekly review.

How this story has evolved

From the week in review: the connections, developments and open questions filed under AI For Science, newest week first.

Week of 7–13 September 2026

Development · Tue 8 Sep, Fri 11 Sep
OpenAI says an internal model with 10,000 sub-agents solved Navier-Stokes; an NYU mathematician says he was pressed to drop an Anthropic-affiliated co-author

On 8 September OpenAI announced "that a multi-agent system, powered and coordinated by an unreleased internal model—that at one point had 10,000 different sub-agents working different parts and variations of the problem—has solved Navier-Stokes", one of the Clay Mathematics Institute's Millennium Prize problems. CNN reports OpenAI said "its model took 88 hours to solve the problem". Fortune puts the compute cost at "about $2 million" on one estimate, with "other reports put the number at 10 times greater still, at $22.5 million".

Development · Tue 8 Sep
DeepMind releases AlphaGenome Atlas: precomputed predictions for 9 billion single-letter DNA changes in a 1-petabyte dataset

Google DeepMind published AlphaGenome Atlas on 8 September, with predictions for the effects of "9 billion single-nucleotide variants — every single-letter change possible" in the human genome, held in "a massive 1-petabyte dataset, more than 30 times larger than the AlphaFold Database".

Open question
Will OpenAI's Navier-Stokes claim be verified, and what happened in the exchange Buckmaster describes?

The Clay Mathematics Institute has issued no determination; CNN reports that "Only one Millenium Prize problem has been officially solved so far". The model is unreleased and OpenAI's announcement page returned HTTP 403, so it was not read for this edition. Fortune derives the roughly $2 million from OpenAI's own briefing statement about compute "at least 1,000 times greater" than a prior about $2,000; the $22.5 million is an outside figure. Buckmaster's account of his exchange with Sebastien Bubeck is his own; Bubeck calls the circulating allegations "false and inflammatory" but his published replies do not address the specific allegation about removing a co-author's name. OpenAI says "we cannot rule out that de-identified data derived from their usage of our products helped improve our models."

Monday, 14 September 2026

Nature Medicine: AI support raised physicians' lung-cancer disease-control sensitivity from 0.72 to 0.87 in a 2,396-patient study beneficial

  • "Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC", published in Nature Medicine on 13 September, reports on I³LUNG (NCT05537922), which enrolled 2,396 patients with stage IIIC–IVB non-small cell lung cancer treated with immunotherapy between September 2012 and October 2023 across six centres in Italy, Greece, Germany, Spain, the USA and Israel.
  • In the usability study, 20 physicians — 10 lung expert oncologists and 10 non-experts — each assessed 10 real-world cases, totalling 200 assessments. With the explainable-AI tool, sensitivity for disease control rate rose from 0.72 (95% CI 0.64–0.80) to 0.87 (95% CI 0.79–0.92), P = 0.0011, and accuracy from 0.57 to 0.65, P = 0.0431, "at the expense of slightly lower specificity". Agreement between experts and non-experts rose from κ = 0.11 to κ = 0.48.
  • Clinical-and-blood-only models "achieved consistent performance across outcomes with area under the curve (AUC) up to 0.77 in the test (TEST) set" and "significantly surpassed PD-L1" and other standard scores in that set. The authors say a prospective validation in more than 2,000 patients is under way.
  • The paper is explicit about limits: performance fell in external validation to an "AUC range: 0.55–0.72", and the benefit of adding CT, pathology and genomics "remains uncertain, not translated in TEST and EXVAL". Physicians adopted correct AI suggestions 74.5% of the time, but experts followed incorrect ones more often than non-experts, 72.2% versus 63.6%.

Harvard-led team releases a fine-tuned physician-level judge and a 9,217-score benchmark for grading medical AI beneficialPreprint

  • "Scaling Clinical Judgment to Evaluate Medical AI" (arXiv 2609.12822, submitted 11 September, announced 14 September) has 17 authors including Thomas A. Buckley, Adam Rodman and Arjun K. Manrai. It introduces PrecepTron, "an LLM fine-tuned for physician-level evaluation of open-ended responses", "trained using low-rank adaptation (LoRA) of a 32-billion-parameter model on a small number of physician examples".
  • The authors release GRAND-ROUNDS, "a new large-scale physician-annotated benchmark of 9,217 scores by 11 physicians across seven studies", and say all code, data and labels are freely available.
  • They report that frontier models used in typical "LLM-as-a-judge" setups "often disagree with physicians and with each other", and say they used PrecepTron to reproduce headline findings from five studies of LLMs for clinical care published in JAMA, Science and Nature Medicine "without new human grading".
  • Preprint, not peer reviewed. Reproducing published findings is not the same as validating clinical safety, and the abstract gives no figure for PrecepTron's agreement with physicians outside the seven studies it was built from.

Sunday, 13 September 2026

UPenn preprint: self-supervised plasma proteomic model predicts 144 diseases across differing protein panels beneficialPreprint

  • The preprint, posted 12 September by Yonghyun Nam, Dokyoon Kim and colleagues at the University of Pennsylvania, reports a self-supervised model built on "53,014 participants in the UK Biobank Pharma Proteomics Project", covering 2,920-protein profiles and a predefined 1,460-protein subset, evaluated across 144 diseases.
  • Reported performance: "median AUC was 0.679 with comprehensive coverage and 0.637 when applied to partial-coverage representations"; retraining only the disease-specific models raised the partial-coverage median AUC to 0.673.
  • The authors report their protein-token risk scores exceeded coefficient-truncated LASSO by a median paired AUC difference of 0.027, and were comparable to LASSO refitted with outcome labels, a median difference of 0.003.
  • This is a preprint and has not been peer reviewed. A median AUC of 0.679 across 144 diseases is a population-level discrimination figure, not a clinical test, and the work is validated inside one cohort.

Preprint reports 99.0% cross-validated sensitivity separating early-stage ovarian cancer from controls in two small cohorts beneficialPreprintSingle source

  • Posted 12 September by Hongyi Zhou, Jean-Luc Chaubard, Benedict Benigno and Jeffrey Skolnick of Georgia Institute of Technology, OmicsIQ LLC and the Ovarian Cancer Institute, the preprint applies boosted decision tree classifiers to blood metabolomic data.
  • Cohorts are "91 serum samples (59 ovarian cancer, 32 healthy controls)" and "83 plasma samples (63 ovarian cancer, 20 healthy controls)". Reported mean cross-validated sensitivity and specificity are 99.0% and 99.8% in serum and 97.7% and 99.7% in plasma.
  • The authors report "239 concordantly altered annotated features spanning lipid, amino-acid, steroid, central-carbon, and redox metabolism", and propose the term "metabolomic Systemotype".
  • The results come from five-fold cross-validation repeated over 50 randomised rounds, not external validation, on fewer than 200 samples in total. Accuracy figures this high on cohorts this small do not establish screening performance in a general population, and the preprint has not been peer reviewed.

Saturday, 12 September 2026

Twenty-five Fields Medallists sign declaration that AI labs' benchmark chasing is "severely misaligned" with mathematics

  • Terence Tao published the declaration on his blog on 11 September; it is signed by 25 Fields Medallists and argues that "solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight".
  • The signatories write that AI solutions are "announced in a rush, leaving no time for a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others", and warn that without that step "AI-conceived ideas would never become fully alive and the crucial human transmission chain between mathematicians would be lost".
  • TechCrunch reports that New York University mathematician Tristan Buckmaster alleged OpenAI pressed him to exclude an Anthropic collaborator from credit on a mathematics problem, and questioned whether OpenAI had drawn on earlier Codex work to produce its own proof.
  • The full text of the declaration is hosted at mathandai.org, which blocked our fetcher, so the quotations above are taken from Tao's own post and TechCrunch. Neither source gives a count of signatories who are also AI-lab collaborators.

OpenAI pulls its $10,000-per-team sponsorship of Caltech's Mathathon after mathematicians' open letter Single source

  • Gizmodo reported on 11 September at 9:05 pm ET that OpenAI research lead Dan Roberts announced by tweet that the company would drop its sponsorship; OpenAI had been supplying $10,000 of the $20,000 in credits available per team, and OpenAI and Anthropic together had pledged $2 million in credits.
  • The withdrawal followed an open letter from current and former Caltech mathematicians saying AI firms have "advanced a campaign of scientific misinformation about the goals of mathematical research" and describing the solutions as having "destructive impacts for the mathematical community".
  • Mathathon organisers told Gizmodo "We do not anticipate that this will affect the event in any substantial way", adding they were "currently in talks with other firms who are willing to provide a similar amount per team". The first round begins 30 October, with each team given 40 hours and $20,000 in tokens.
  • Gizmodo is the only outlet we could open carrying the dollar figures; OpenAI did not give a statement in the piece beyond Roberts's post.

Jeff Dean's Discovery Loop seeking new funding at about $50bn, five times its valuation of a few weeks ago Single source

  • Business Insider reported on 11 September that Discovery Loop is seeking funding at a valuation of roughly $50 billion, up from the approximately $10 billion valuation at which it was seeking $1 billion just weeks earlier.
  • The company was founded by former Google chief scientist Jeff Dean with Sanjay Ghemawat, Quoc Le and Oriol Vinyals, and says it aims to "use AI to accelerate scientific and engineering research through the parallel execution of thousands of experiments".
  • Its initial funding was led by Radical Ventures and Khosla Ventures with Lightspeed, Kleiner Perkins and Doerr Capital participating; Alphabet is a founding investor and cloud partner.
  • A Discovery Loop spokesperson declined to comment and Dean did not respond. There is no confirmation the financing will close at that valuation, and the company has no disclosed commercial product.

Mayo Clinic study: AI reading of routine slides links tumour spatial pattern to 71% higher pancreatic cancer recurrence risk beneficial

  • The study, published in Clinical Cancer Research on 11 September as "Spatial Configuration of Pancreatic Cancer Is Associated with Disease Recurrence after Neoadjuvant Therapy and Curative-Intent Resection" (DOI 10.1158/1078-0432.ccr-25-4968), analysed tissue from 203 patients with pancreatic ductal adenocarcinoma.
  • The AI measured "tissue shape, fragmentation and the degree to which the cancer and stroma were intermixed" on standard pathology slides. "High-risk patients had a 71% higher adjusted risk of recurrence" in one model and "more than twice the adjusted risk" in another, while "the amount of residual cancer alone did not reliably separate patients at higher and lower risk".
  • High-risk spatial patterns "contained fewer immune cells within the cancer itself, with immune cells tending to collect around the tumor instead of entering it" — a mechanism, not just a correlation, and one that runs on slides hospitals already produce.
  • Dr Ryan Carr said "the results are promising but need to be confirmed in prospective studies before this approach could be used to inform clinical decision-making". Mayo Clinic's own newsroom page blocked our fetcher, so the figures above are as Medical Xpress reports them.