Tuesday, 22 September 2026 / transcript

Transcript — Tue 22 Sep

Episode cover
0:00 / 16:11
The AI Edge · Maya & Alex · 16:11 · read the transcript · subscribe · open in Spotify

Maya and Alex are AI voices. Each part of the conversation below comes from one item in the written edition — linked above it — and is checked automatically before publishing: every number must appear in that item, every caveat the edition raises must be said aloud, the source must be named, and speculative or hyped language is rejected.

Intro
MayaIt's Tuesday, September 22nd. This is The AI Edge, presented by Epilogue. I'm Maya.
AlexAnd I'm Alex. Both of our voices are AI-generated. Where we could only read a relay rather than the original, we'll say so.
MayaThis is the last 24 hours at the frontier of AI. What got built, what the research found, and how the technology is being used, for good and for harm.
AlexThree things lead today. OpenAI says an internal model it began training on August 28th has resolved more than 100 long-standing open problems across most areas of mathematics.
MayaAlibaba went the other way on pace. It says the Qwen 4.5 and Qwen 5 series are projected to scale up to 5 to 10 trillion parameters, and that Alibaba Cloud will operate more than 20 gigawatts of data centre capacity by 2032.
AlexAnd the buildout ran into policy. Texas ordered its environmental regulator to issue no data centre permits until grid and water audits are complete, and California's governor signed seven data centre laws.
MayaLet's start with the mathematics claim. TechCrunch reports OpenAI said an internal model has resolved more than 100 additional open problems across most areas of mathematics, after its claimed solution to the Navier-Stokes Millennium Prize problem.
AlexThat's a company claim, and it hasn't been independently verified. There's no list of the problems, no proofs, no referee reports published alongside it.
MayaAnd only one outlet. We couldn't open OpenAI's own post at all, so everything we have is the way TechCrunch reports it.
AlexOpenAI also announced an advisory group on mathematics and AI, hosted at the Institute for Advanced Study in Princeton. Members aren't paid, and per TechCrunch the group will not be responsible for advising OpenAI on how to pace its internal progress.
MayaWhich is the interesting limit. Of the initial members, TechCrunch says only one also signed the open letter from 25 Fields Medal winners objecting to how AI labs have behaved in mathematics.
AlexAlso on Monday, OpenAI published a set of frontier-safety proposals. CNBC quotes the post saying fully autonomous recursive self-improvement is not happening today, and that we should not pursue it unless and until it can be done safely.
MayaWhat's the worry they name?
AlexCNBC quotes them saying that done without appropriate care and caution, it could result in humans losing practical control over AI development, unable to provide oversight on research processes they no longer understand.
MayaThe post also cites the Hugging Face agent hack, which CNBC notes did not involve that technique, as a preview of risks that could become much more severe without robust safeguards and alignment.
AlexOne caveat. The OpenAI post itself wouldn't open for us either, so this is a single source. Every quotation is CNBC's rendering of it, not text we read on OpenAI's site.
MayaNow hold that next to Alibaba, on the same day. Alibaba Cloud's release says that over a month of fully automated runs, Qwen3.8-Max completed 33 iterative cycles, and that the updated model boosted its Artificial Analysis score from 40 to 45.
AlexSo one lab says don't pursue this yet, and the other advertises it as a feature.
MayaThat's the shape of it. Alibaba Cloud also describes a chip design experiment where the model ran more than 60 hours of self-improvement, made more than 10,000 EDA tool calls, and reduced chip area by 42% with what the release calls zero compromise in performance.
AlexEvery one of those is a company claim, and none is independently verified. The release doesn't say what human oversight those automated runs had, what the 42% was measured against, or whether the improved model has been deployed.
MayaThere was also a model release. xAI's page lists Grok 4.7 xHigh at $2 per million input tokens and $6 per million output tokens, against $4 and $20 for GPT-5.6 Sol Max.
AlexAnd how does it actually score?
MayaOn xAI's own numbers, 46.3% on CursorBench 4.0, against 40.4% for Grok 4.6. Those are self-published benchmarks and they're not independently verified.
AlexThe Decoder cites the Artificial Analysis index putting Grok 4.7 at 46, against 53 each for Claude Fable 5.1 and GPT-6. And the two sources flatly disagree on agentic coding.
MayaThey do. xAI's page shows 38.0% on Terminal-Bench 4.0. The Decoder reports Artificial Analysis measuring 26%. Neither source explains the gap, and we didn't reconcile it.
AlexTo the research, and it lands on the same theme. A paper on arXiv from Google Cloud AI Research and university co-authors puts guardrails on an agent improving its own harness, and reports gains of up to 14.1 points on the split it evolves against.
MayaAnd away from those benchmarks?
AlexUp to 4.7 points on the five out-of-distribution benchmarks, while running on 30% fewer policy tokens than the unregularised version. That gap between 14.1 and 4.7 is the paper's own central caveat.
MayaIt's a preprint, not peer reviewed, and the results are the authors' own, so treat them as a company claim. Worth saying the model itself is frozen here. Only the scaffolding around it evolves.
AlexIn a paper on arXiv, a political scientist at Purdue ran a preregistered experiment of 7,500 multi-turn conversations across six AI assistants, randomly assigning the user's political identity.
MayaWhat did it find?
AlexThe paper reports that on abortion, GPT engages and mirrors every user, Gemma refuses everyone, Claude answers strongly conservative users 35 percent of the time and almost no one else, and Grok accommodates conservatives only.
MayaOn climate change and Nazism the paper reports five systems hold firm for every user. It also says comparing two Grok releases shows the regime changing between versions in a way current audits miss.
AlexCaveats matter here. It's a preprint, it's a single source, and the answers were scored by AI judges rather than human coders. Those model behaviours are the paper's characterisations, and we didn't reproduce them.
MayaTo security. Reuters reports the Chinese startup Z.ai disabled some features of its flagship AI coding assistant after users reported it was uploading entire local code repositories onto overseas cloud servers without their consent.
AlexHow did that happen?
MayaZ.ai said it came from a codebase indexing feature that was on by default, and that it patched the vulnerability. On Monday it also open-sourced the assistant and pledged to make the product more transparent.
AlexThere's a complication. One firm said on Friday that six of its coding workspaces were uploaded without consent, including complete source code, database passwords and employees' personal information. On Monday it retracted that, saying it had wrong evidence.
MayaSo the scale is unresolved. Reuters also reports users said the deleted data was encrypted with a key held only by Z.ai, so they couldn't verify deletion. Reuters calls this a rare public disclosure of a security breach by a Chinese AI lab.
AlexHere's one about filed financial documents. A benchmark posted to arXiv measures how reliably a coding agent alters one dollar amount, date or address in a real filed financial document, from a single sentence of intent.
MayaAnd the numbers?
AlexAcross 1,750 cells, the paper reports 1,419, or 81.1%, satisfy the verifier, and 808, or 46.2%, also survive every stricter filter. The cheapest verified forgery costs 2.4 cents, and no model refused.
MayaThe control is the part I'd hold onto. A deterministic script with no model in it solves 98 of the 125 documents. The agents solve 124, and none that the script solves alone.
AlexSo agency buys coverage, not a brand new capability. It's a preprint, it's a single source, and the authors list a company that sells detection, which is a commercial interest in the finding. It documents no real-world fraud using these methods.
MayaAnd a useful counterweight. The Register reports that of 225 vulnerabilities linked to Anthropic or Project Glasswing and tracked by a VulnCheck researcher, just one has confirmed exploitation against real targets.
AlexThe researcher is quoted saying the main thing the data highlights is that what Anthropic is discovering and disclosing is fairly limited in impact.
MayaThe same piece cites 1Password research on 6,080 patches from two frontier models. They fully resolved the vulnerability just 26 percent of the time, and about 54 percent either failed to fix it, introduced a new one, or did both.
AlexThat's one researcher's tracking reported by a single source. And it's a snapshot. Exploitation can lag disclosure by months.
AlexOn to geopolitics, and this is an update to something we covered yesterday. In CNBC's published transcript, Treasury Secretary Scott Bessent said the two countries have now formalised something called the USA-China AI dialogues, and have agreed to meet again in two months in Shenzhen. He put that timing loosely.
MayaWhat would the mechanism actually do?
AlexHe said they want to open a communications line, an incident line, so there's constant communication especially in the event of some kind of incident. And that both sides want to start discussing protocols on what the leading AI dangers are.
MayaHe named uncontrollable agents, non-state actors, and cyber non-state actors in bioweapons. He also said the talks with the Chinese vice premier ran about 12 hours.
AlexWhat's new since yesterday is the name, the venue and the cadence. Nothing here is a signed agreement, there's no published Chinese confirmation of the terms, and he named no protocol that's actually been agreed.
MayaNow a result on the beneficial side. Nature Medicine published a model called EAGLE that detects esophageal cancer and precancerous lesions from chest noncontrast CT, a task the paper describes as historically considered impossible.
AlexHow much validation is behind it?
MayaIt was trained on 6,813 patients from two centres and validated across 12 centres in three countries involving 80,612 patients. On external test cohorts, 98.5% specificity with 90.0% sensitivity for cancer.
AlexAnd the appeal is that it reads scans people are already getting, including inside lung cancer screening programmes.
MayaThe weak point is precancerous lesions, at 52.5% sensitivity in those external cohorts. And the paper registers no outcome trial showing the model changes mortality. Detection accuracy is not the same thing as a patient living longer.
AlexTo policy, and this came from the same CNBC interview. Bessent said he agrees with the MIT professor Daniel Huttenlocher that it is humans who are responsible, not the AI, and that the Hugging Face incident is the responsibility of the OpenAI management, not a bunch of agents.
MayaAnd on the labs asking to be indemnified?
AlexHe was blunt. He said the labs also asked to take the liability off their hands, and, quote, we will not do that. He added that these labs need to take responsibility for themselves, and they can slow down any time they want to.
MayaThat's a cabinet secretary naming a specific company as accountable for a specific incident, on the same day OpenAI published its own safety proposals.
AlexWorth keeping in proportion, though. These are remarks in a television interview, not a rulemaking, an enforcement action or a bill. And Treasury isn't the agency that would set AI liability.
MayaOn the infrastructure side, Alibaba unveiled a new accelerator, the Zhenwu V900. Alibaba Cloud's release says it delivers three times the performance of its predecessor, with 216 gigabytes of GPU memory, and mass production in the first quarter of 2027.
AlexAnd the capacity target?
MayaThe chief executive is quoted saying that by 2032, the global data centre capacity operated by Alibaba Cloud will surpass 20 gigawatts. The company also says its upgraded server can support a cluster of up to 500,000 cards.
AlexThe performance, memory and bandwidth figures are a company claim, unverified, and no benchmark results against Nvidia parts were published. The chip isn't in mass production until the first quarter of 2027, and the capacity number is a target for 2032.
AlexAnd the other side of the buildout. The Texas Tribune reports Governor Greg Abbott ordered the state environmental regulator to halt all environmental permit approvals for data centre projects until grid and water audits are complete.
MayaHe's quoted saying Texans must come first, that data centres must pay their own way and protect the grid and water, and that until they do, the commission will issue no permits sought by data centre projects.
AlexOne number stands out. Only 28% of data centres responded to a state-mandated water usage survey, which prompted him to direct penalties for non-compliance.
MayaThis is an update rather than a new posture. The audits and a grid connection moratorium were ordered in August, and only the permit halt is new. It's a single source, and the Tribune doesn't say how many pending projects are affected.
MayaLast one, and it's the hardest. MIT Technology Review cross-referenced nearly 4,000 locations where human remains were found against nearly 600 border surveillance towers identified by the Electronic Frontier Foundation.
AlexWhat did they find?
MayaMore than 1,050 people who died within range of those towers between 2015 and early 2026, inside the advertised range of nearly two-thirds of the towers analysed. Their terrain analysis found some towers can see as little as 10% of their advertised area.
AlexAnd the programme is growing. They report the towers now number 803, with a 2023 government estimate of $6.2 billion over their lifespan, and plans to spend $1 billion for 1,497 more by 2034.
MayaCustoms and Border Protection says the autonomous towers use artificial intelligence to detect and classify people, vehicles and animals. Anduril said a death nearby does not mean a tower missed a detection.
AlexAnd this is one outlet's own analysis. Proximity to a tower is not evidence that a tower caused or could have prevented a death, and the publication doesn't claim that it is.
Outro
MayaThat's The AI Edge for today.
AlexThe full edition, with a link to every source behind every claim, is on the site.
MayaListen in tomorrow for the next edition.