Topics / topic

Google DeepMind

6 items across 4 editions · appeared in the last 4 editions in a row. First seen Sat 12 Sep, last seen Tue 15 Sep. Traced across 1 weekly review.

How this story has evolved

From the week in review: the connections, developments and open questions filed under Google DeepMind, newest week first.

Week of 7–13 September 2026

Connection
Two government agencies, two frontier labs and Beijing all spoke about model extraction within three days

The joint advisory came on 8 September from CISA, NSA and FBI; Google's threat group published the same day; Anthropic's report followed on 10 September; Beijing responded on 9 September; and Amodei's essay of 12 September asks governments to "Crack down on unauthorized distillation by companies in authoritarian countries" and to "Do not sell powerful AI chips or semiconductor manufacturing equipment to China".

Development · Tue 8 Sep, Wed 9 Sep, Thu 10 Sep
Anthropic names seven Chinese labs over illicit distillation and Google reports campaigns exceeding 100 million prompts, two days after a joint US advisory named six firms

On 8 September CISA, NSA and FBI issued joint advisory AA26-251A, "China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies", naming DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI. The advisory says the firms "extracted billions of tokens across millions of exchanges/requests from U.S. frontier AI models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024", and says DeepSeek's publicly stated $5.6 million training cost excludes data "acquired through extensive malicious distillation".

Development · Tue 8 Sep
DeepMind releases AlphaGenome Atlas: precomputed predictions for 9 billion single-letter DNA changes in a 1-petabyte dataset

Google DeepMind published AlphaGenome Atlas on 8 September, with predictions for the effects of "9 billion single-nucleotide variants — every single-letter change possible" in the human genome, held in "a massive 1-petabyte dataset, more than 30 times larger than the AlphaFold Database".

Open question
Will any company other than Anthropic put an embedded-evaluator commitment in writing, and with which evaluator?

Anthropic's is the only commitment published as a document, and it names no start date. OpenAI's position is a policy post plus Altman's statement that "We'll have more to share soon". Musk's and Hassabis's statements are brief endorsements rather than commitments, and Sunak states he is a senior adviser at Anthropic. No source has named which organisation would embed reviewers at OpenAI, Google DeepMind, Microsoft or xAI, on what terms, or with what right to publish.

Tuesday, 15 September 2026

Google DeepMind AGI safety researcher publicises resignation, saying AI "has the potential to kill us all" Single source

  • The News International, reporting on 15 September, quotes Bilal Chughtai's post on X: "I recently resigned from Google DeepMind, where I worked on AGI safety and alignment research. I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome."
  • The report says Chughtai worked as a research engineer on AGI safety and alignment research and left DeepMind in July 2026, posting the warning this week.
  • The article carries no response from Google or Google DeepMind. The post is a personal statement: it is not accompanied by evaluation data, internal documents or any specific capability claim.

Google Research and CMU harness scores 71.0% on research-level TCS-Bench, solves 218 of 222 Codeforces problems PreprintCompany claim

  • arXiv:2609.15983, "Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science" by Honghao Lin, David P. Woodruff, Yuan Deng, Jieming Mao, Song Zuo and Vahab Mirrokni, submitted 14 September 2026, reports: "On TCS-Bench, a benchmark of research-level theorem-proving tasks drawn from papers published at FOCS, STOC, and SODA, Colosseum achieves 71.0% accuracy using Gemini 3.1 Pro and Gemini 3.7 Flash."
  • The paper reports that in a separate Codeforces evaluation using Gemini 3.1 Pro, "the proof-oriented pipeline with execution feedback solves 218 of 222 problems", and that with Gemini 3.1 Pro the authors "obtain several new results that address open problems arising from papers published at top venues such as FOCS and JMLR".
  • The authors state the workflow "has also been integrated into Google Antigravity's Teamwork framework as the Long Proof pattern", so the harness is already shipping inside a Google product.
  • The paper is a preprint by authors at the company whose models it evaluates, the abstract does not name the open problems said to be resolved, and the claimed new results have not been independently checked.

Monday, 14 September 2026

The Information: Google, Anthropic and OpenAI have met regularly since July about an industry AI standards body Single sourceUpdate

  • PYMNTS, citing a report published by The Information on 13 September, says representatives from Google, Anthropic and OpenAI "have been regularly meeting since July about the proposal for a standards body" covering testing and auditing of frontier models.
  • According to PYMNTS, OpenAI chief executive Sam Altman has voiced support at a company town hall for "a testing and auditing organization for the industry" but believes the major labs should set standards without the backing of the US government, while Amodei's framework allows for voluntary corporate standards alongside government regulation.
  • PYMNTS says the discussions follow an essay published by Demis Hassabis in July 2026 proposing a self-regulatory body modelled on the Financial Industry Regulatory Authority.
  • The Information's article is paywalled and was not read directly; these facts come from PYMNTS' account of it. Nothing has been finalised, no body has been chartered or named, and none of the three companies has published terms.

Expert re-grading finds 238 of 250 failed physics-benchmark answers were benchmark or grader errors, not model errors mixedPreprint

  • "How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks" (arXiv 2609.13009, submitted 11 September, announced in the 14 September listing) has 51 authors; the HTML version lists Yale University and Jump Trading Group among the affiliations. Physics faculty and their graduate researchers audited text-only, closed-ended questions in their own subfields across six benchmarks.
  • The audit covered 502 questions. Of the 250 rejected answers sent for review, 143 (57.20%) were classified as benchmark errors — a defective problem statement or reference solution — 95 (38.00%) as grader errors and 12 (4.80%) as genuine model errors; 238 of the 250, or 95.20%, were benchmark or grader errors.
  • The abstract reports GPT-5.6-Sol's measured mean@4 rising from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, with corrected pass@4 reaching 94.4% on the 54 retained CritPt challenges. The authors write that "current benchmarks substantially understate frontier models' ability to solve well-posed physics problems".
  • This is a preprint and has not been peer reviewed. Corrected scores are computed on retained subsets after flawed questions were repaired or excluded, so they are not like-for-like with the original figures, and the audit covers only text-only closed-ended questions with verifiable answers.

Sunday, 13 September 2026

Altman, Musk, Hassabis and Sunak back Amodei's pacing proposal; OpenAI says it will adopt embedded evaluators Company claim

  • CNBC reports Altman posted on X that pacing has been a "primary topic" of discussion at OpenAI in recent weeks, adding: "Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon." Musk wrote "Dario is right." CNBC calls it "an unusual show of agreement among three fierce rivals".
  • The Tribune, published 13 September at 7:02 AM IST, quotes Google DeepMind's Demis Hassabis: "Dario's essay points towards the right path forward. The details need working through, but the direction is correct for meeting this critical moment." Former UK prime minister Rishi Sunak, who states he is a senior adviser at Anthropic, also endorsed the proposal.
  • CNBC notes OpenAI chief scientist Jakub Pachocki published a blog post earlier this month saying no AI company has "solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer", and that he expects voluntary slowdowns to become "commonplace until shared safety bars are established".
  • These are statements of intent on social media, not published commitments. OpenAI has not said when evaluators would be embedded or on what terms, and no company besides Anthropic has published an access agreement.

Saturday, 12 September 2026

Jeff Dean's Discovery Loop seeking new funding at about $50bn, five times its valuation of a few weeks ago Single source

  • Business Insider reported on 11 September that Discovery Loop is seeking funding at a valuation of roughly $50 billion, up from the approximately $10 billion valuation at which it was seeking $1 billion just weeks earlier.
  • The company was founded by former Google chief scientist Jeff Dean with Sanjay Ghemawat, Quoc Le and Oriol Vinyals, and says it aims to "use AI to accelerate scientific and engineering research through the parallel execution of thousands of experiments".
  • Its initial funding was led by Radical Ventures and Khosla Ventures with Lightspeed, Kleiner Perkins and Doerr Capital participating; Alphabet is a founding investor and cloud partner.
  • A Discovery Loop spokesperson declined to comment and Dean did not respond. There is no confirmation the financing will close at that valuation, and the company has no disclosed commercial product.