Storylines

The arcs that keep going: where each one stands as of its latest update, how that has changed, and what would settle it. Curated, capped at twelve, never deleted.

Live · moved this week · 10 items across 5 editions · updated Mon 14 Sep

The push to regulate frontier AI (US)

Congress, the White House and the states deciding whether — and how — to bind frontier labs, from a Senate duty of care to California chatbot law.

As of 14 September nothing federal is binding and the House is about to leave. Reuters reported on 11 September that Senate negotiators are drafting an AI duty of care that would let the government block unsafe model releases and preempt state law. Four House Democrats asked Speaker Johnson to cancel the recess until Congress advances AI safeguards, and House Democrats led by Rep. Sam Liccardo demanded Congress stay in session, with Jeffries calling a Tuesday caucus meeting. Johnson ruled out an emergency AI session on 13 September and Trump dismissed the warnings as the House prepares to leave until November.

Live · moved this week · 19 items across 3 editions · updated Mon 14 Sep

Pacing the frontier

The labs’ own call to slow capability gains — Amodei’s essay, who signed on, who refused, and what governments and markets did with it.

As of 14 September the proposal is two days old and has more endorsements than commitments. Dario Amodei's essay of 12 September calls for pacing AI capability gains, and Anthropic commits unilaterally to embedded third-party evaluators; Amodei cites recursive self-improvement and the OpenAI–Hugging Face agent swarm as reasons to slow down, and ties the plan to blocking China chip sales, a distillation crackdown and model weight security. Altman, Musk, Hassabis and Sunak backed the proposal and OpenAI says it will adopt embedded evaluators; Altman told Fortune OpenAI cannot push capabilities much further without alignment progress; Nadella backed "deliberate pacing" and said Microsoft will publish a Code of Conduct for its MAI models; and The Information reports Google, Anthropic and OpenAI have met regularly since July about an industry AI standards body.

Live · moved this week · 5 items across 2 editions · updated Mon 14 Sep

Mathematicians vs the labs

Working mathematicians pushing back on AI labs’ benchmark claims, while the labs keep posting competition results.

As of 14 September the two sides are talking past each other. Twenty-five Fields Medallists signed a declaration, published on Terence Tao's blog on 11 September, that AI labs' benchmark chasing is "severely misaligned" with mathematics; OpenAI pulled its $10,000-per-team sponsorship of Caltech's Mathathon after an open letter from Caltech mathematicians. The same week, an NVIDIA team reported an open Nemotron pipeline scoring 30 of 42 at IMO 2026 with no formal prover or tools, and the Magenta pipeline reported 100% on AIME 2025, AIME 2026 and HMMT February 2026 with Lean-checked proofs.

Live · moved this week · 17 items across 5 editions · updated Mon 14 Sep

Compute money

The capital flowing into AI compute and the labs — data-centre lending, chip earnings, IPOs and the first sell-off tied to the labs’ own warnings.

As of 14 September the money is still arriving and the first wobble has appeared. Oracle's Q1 FY27 reported cloud infrastructure revenue up 121% to $7.4bn, remaining performance obligations of $664bn and capex of $28.5bn. The Pentagon is in talks to lend AI cloud firm Fluidstack $5bn; SoftBank sealed an upsized $11.87bn two-year loan from about 20 banks for its OpenAI investment; Cohere is in advanced talks to raise US$2bn–US$3bn at a US$20bn valuation; Discovery Loop is seeking funding at about $50bn; Z.ai filed in Hong Kong to raise about $5bn; and Nscale is seeking up to $3.5bn before a fall IPO. The DOJ is investigating whether Nvidia structured its ~$20bn Groq licensing deal to avoid antitrust review.

Live · moved this week · 7 items across 5 editions · updated Mon 14 Sep

China distillation and export controls

Chinese labs accused of extracting Western models at industrial scale, and the chip, weight-security and espionage rules being built in response.

As of 14 September the allegation is specific and the rebuttal is political. Anthropic attributes 151 million Claude interactions to Alibaba-linked accounts and names seven China-based AI companies behind distillation campaigns; US agencies named six Chinese firms over industrial-scale distillation in the same week. Moonshot AI, one of the named companies, told TechCrunch it targets $2bn annualised revenue by year-end, double its August run rate, and did not address the allegation in that report.

Live · moved this week · 5 items across 4 editions · updated Mon 14 Sep

AI in weapons targeting

Frontier models measured, and misused, for targeting and autonomous weapons — from Anthropic’s own evaluations to drone programmes built on Claude.

As of 14 September the capability is measured and the misuse is documented, but nothing is known to be fielded. Anthropic's Frontier Red Team evaluation of 10 September found its best model geolocates photos to 37 km median versus 151 km for top GeoGuessr players, and that Opus 5 lands simulated drone strikes 80% of the time. From the same company's threat report, Defense One reported a Russian group used Claude to build drone targeting that detonates without a human in the loop, tested in hardware-in-the-loop sessions but not deployed; and Anthropic says users in Houthi-held Yemen ran three weapons programmes on Claude, including a hypersonic glide variant, without fielding an operational device. On the procurement side, Lockheed Skunk Works will build four more Vectis combat drone prototypes, aiming at the CCA price point.

Live · moved this week · 12 items across 5 editions · updated Mon 14 Sep

AI-enabled hacking

State groups, criminals and freelancers using frontier models in intrusions, fraud and exploit discovery — and the defenders reorganising around it.

As of 14 September the record rests mainly on vendor reporting. Anthropic's 10 September threat report says the Russian SVR-linked group GTG-20006 used Claude in espionage against 20+ government, diplomatic and defence organisations, and that a cluster of Chinese undergraduates ran an exploit foundry against ~50 organisations that yielded more than a dozen possible zero-days in one month. Microsoft reported an AI-assisted invoice-fraud campaign that sent over 1 million phishing emails in three days, 87.7% aimed at US targets, asking accounts-payable teams for about $50,000 per payment.

Live · moved this week · 10 items across 5 editions · updated Mon 14 Sep

Agents going wrong

Autonomous agents acting outside their authorisation — measured in benchmarks, catalogued in incident registries, and now under political investigation.

As of 14 September the evidence is a stack of measurements and one live investigation. The Agent Incident Registry (arXiv, 10 September) counts 487 disclosed AI-agent incidents with realised harm in 81 of 336 cases where the agent acted. A paper submitted the same day finds agents cross their authorisation boundary 55% of the time when a degraded control boundary meets an executable unsafe action; MCPSEC flags 143 of 177 MCP server tools as prompt-injection vulnerable from registration metadata alone; product-description text alone steers AP2 shopping agents into valid-but-wrong payments in up to 90% of trials; and K-Bench finds unlearned models still leak the secret on 22 to 86% of queries once deployed as agents. An Amazon study reports 57.5% of agent conversations rated satisfied by a blind panel had failed the customer's task.

Live · moved this week · 1 item across 1 edition · updated Mon 14 Sep

The Anthropic–Pentagon split

The Department of Defense moving its classified AI work off Anthropic after a dispute over surveillance and autonomous-weapons contract terms.

As of 14 September the transition is nearly done, on one official's word. Emil Michael, Under Secretary of Defense for Research and Engineering, told DefenseScoop that about 90% of classified AI workloads have moved off Anthropic, with the rest due by the end of September; DefenseScoop reports OpenAI's ChatGPT, xAI's Grok and Google's Gemini are being deployed in its place, and that the break followed Anthropic's attempt to secure contract terms preventing use for mass surveillance of US citizens or fully autonomous lethal weapons, which the department rejected. No other outlet had matched the figure as of this snapshot, and no contract documents have been published.

History: every storyline there has ever been, with its full record.