Storylines / Live · opened Mon 14 Sep · moved this week · 5 items · 1 update

Mathematicians vs the labs

Working mathematicians pushing back on AI labs’ benchmark claims, while the labs keep posting competition results.

What would settle it

Does the dispute change how labs report mathematical results, or how mathematicians engage with them? Settled by: a lab adopting the declaration’s write-up norms, a joint statement, or the Mathathon proceeding without lab sponsorship.

Where this stands as of Monday, 14 September 2026

As of 14 September the two sides are talking past each other. Twenty-five Fields Medallists signed a declaration, published on Terence Tao's blog on 11 September, that AI labs' benchmark chasing is "severely misaligned" with mathematics; OpenAI pulled its $10,000-per-team sponsorship of Caltech's Mathathon after an open letter from Caltech mathematicians. The same week, an NVIDIA team reported an open Nemotron pipeline scoring 30 of 42 at IMO 2026 with no formal prover or tools, and the Magenta pipeline reported 100% on AIME 2025, AIME 2026 and HMMT February 2026 with Lean-checked proofs.

Timeline

Every item filed under this storyline, newest first, with its sources.

Tuesday, 15 September 2026

Tue 15 Sep · Research & papers

Google Research and CMU harness scores 71.0% on research-level TCS-Bench, solves 218 of 222 Codeforces problems PreprintCompany claim

  • arXiv:2609.15983, "Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science" by Honghao Lin, David P. Woodruff, Yuan Deng, Jieming Mao, Song Zuo and Vahab Mirrokni, submitted 14 September 2026, reports: "On TCS-Bench, a benchmark of research-level theorem-proving tasks drawn from papers published at FOCS, STOC, and SODA, Colosseum achieves 71.0% accuracy using Gemini 3.1 Pro and Gemini 3.7 Flash."
  • The paper reports that in a separate Codeforces evaluation using Gemini 3.1 Pro, "the proof-oriented pipeline with execution feedback solves 218 of 222 problems", and that with Gemini 3.1 Pro the authors "obtain several new results that address open problems arising from papers published at top venues such as FOCS and JMLR".
  • The authors state the workflow "has also been integrated into Google Antigravity's Teamwork framework as the Long Proof pattern", so the harness is already shipping inside a Google product.
  • The paper is a preprint by authors at the company whose models it evaluates, the abstract does not name the open problems said to be resolved, and the claimed new results have not been independently checked.

Saturday, 12 September 2026

Sat 12 Sep · Research & papers

Magenta pipeline reports 100% on AIME 2025, AIME 2026 and HMMT February 2026 with Lean-checked proofs Preprint

  • The paper, posted to arXiv on 10 September 2026, describes "a training-free agentic pipeline that, given only a natural-language problem, produces an answer, expresses it as a Lean 4 statement, and constructs a machine-checked proof", and reports "100% accuracy across all evaluated olympiad benchmarks, including AIME 2025, AIME 2026, and HMMT February 2026".
  • The abstract states that when paired with the open-weight K2-Horizon-7B reasoner "it solves all six IMO 2026 problems".
  • The design puts a statement judge in front of the prover to check that the Lean formalisation preserves the original problem, and an error-attribution judge that routes failures either to mathematical re-derivation or to local Lean repair — the guard against a proof that is machine-checked but of the wrong statement.
  • This is a preprint with no peer review, and the benchmark figures are the authors' own runs. The abstract does not report compute cost or the number of attempts per problem.
Sat 12 Sep · Research & papers

NVIDIA team reports an open Nemotron pipeline scoring 30 of 42 at IMO 2026 with no formal prover or tools PreprintCompany claim

  • The paper, posted to arXiv on 9 September 2026 and announced in the 11 September listing, reports that the system "scored 30 out of 42 points at IMO 2026, reaching the gold-medal threshold".
  • The abstract states the pipeline "operates entirely in natural language, with no formal prover, external tools, or internet access", using three Nemotron 3 Ultra checkpoints — the general-availability model and two post-trained specialists — in "an iterative search that generates, verifies, and refines candidate proofs", with a separate high-compute stage selecting each submission.
  • The authors say they release both post-trained checkpoints plus the training data, training and inference code, the submitted solutions, and "Nemotron-IMO-Bench, a new benchmark of 200 novel olympiad-level problems".
  • The arXiv abstract page carries no affiliation block; Hugging Face lists the institution as NVIDIA. The IMO score is the authors' own report of their own submission and is not peer reviewed.
Sat 12 Sep · Frontier models & labs

OpenAI pulls its $10,000-per-team sponsorship of Caltech's Mathathon after mathematicians' open letter Single source

  • Gizmodo reported on 11 September at 9:05 pm ET that OpenAI research lead Dan Roberts announced by tweet that the company would drop its sponsorship; OpenAI had been supplying $10,000 of the $20,000 in credits available per team, and OpenAI and Anthropic together had pledged $2 million in credits.
  • The withdrawal followed an open letter from current and former Caltech mathematicians saying AI firms have "advanced a campaign of scientific misinformation about the goals of mathematical research" and describing the solutions as having "destructive impacts for the mathematical community".
  • Mathathon organisers told Gizmodo "We do not anticipate that this will affect the event in any substantial way", adding they were "currently in talks with other firms who are willing to provide a similar amount per team". The first round begins 30 October, with each team given 40 hours and $20,000 in tokens.
  • Gizmodo is the only outlet we could open carrying the dollar figures; OpenAI did not give a statement in the piece beyond Roberts's post.
Sat 12 Sep · Frontier models & labs

Twenty-five Fields Medallists sign declaration that AI labs' benchmark chasing is "severely misaligned" with mathematics

  • Terence Tao published the declaration on his blog on 11 September; it is signed by 25 Fields Medallists and argues that "solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight".
  • The signatories write that AI solutions are "announced in a rush, leaving no time for a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others", and warn that without that step "AI-conceived ideas would never become fully alive and the crucial human transmission chain between mathematicians would be lost".
  • TechCrunch reports that New York University mathematician Tristan Buckmaster alleged OpenAI pressed him to exclude an Anthropic collaborator from credit on a mathematics problem, and questioned whether OpenAI had drawn on earlier Codex work to produce its own proof.
  • The full text of the declaration is hosted at mathandai.org, which blocked our fetcher, so the quotations above are taken from Tao's own post and TechCrunch. Neither source gives a count of signatories who are also AI-lab collaborators.

Tracked figures

30 of 42
IMO 2026 score reported for the open Nemotron pipeline Fri 11 Sep arXiv
25
Fields Medallists signing the declaration Fri 11 Sep Terence Tao

Open questions

Will the Mathathon go ahead, and with whose sponsorship?

What would settle it
The organisers naming replacement sponsors or cancelling; the first round begins 30 October.
Asked
Mon 14 Sep