Monday, 5 October 2026 / transcript

Transcript — Mon 5 Oct

Episode cover
0:00 / 14:41
The AI Edge · Maya & Alex · 14:41 · read the transcript · subscribe

Maya and Alex are AI voices. Each part of the conversation below comes from one item in the written edition — linked above it — and is checked automatically before publishing: every number must appear in that item, every caveat the edition raises must be said aloud, the source must be named, and speculative or hyped language is rejected.

Intro
MayaIt's Monday, October 5th, and this is The AI Edge, presented by Epilogue.
AlexEpilogue is an AI venture studio and consultancy in Toronto, building products where the answer has to be right. Find out more at epiloguelabs.com.
MayaI'm Maya.
AlexAnd I'm Alex.
MayaHere's what moved at the frontier of AI since yesterday morning: the advances, the research, and the uses for good and for harm, with every claim linked to its source on the site.
AlexWhat's at the top today?
MayaFirst, senior executives from Anthropic, OpenAI, Google and Meta are testifying under oath today before a rare New York City Council committee where all 51 members are expected to attend.
AlexSecond, the Internet Watch Foundation says it assessed 6,310 child sexual abuse images generated by AI in the first half of this year, which is 40% more than the 4,512 it recorded across all of 2025.
MayaAnd third, MLCommons published the first version of its jailbreak benchmark, and found the unsafe-response rate across eight open-weight systems rose from 11.08% at baseline to 18.65% under attack.
Transition
AlexLet's start with the labs.
MayaBenzinga reports that in an interview set to appear in Monday's debut edition of Politico's Decoded, Sam Altman said OpenAI and Anthropic have, in his words, a lot of daylight between their policy positions.
AlexWhat did he actually say?
MayaBenzinga quotes him saying: we believe that the world should accept some bad things happening for the benefits of this technology and people having agency.
AlexAnd on regulation?
MayaBenzinga reports he rejected concerns about regulatory capture, and argued that concentrating powerful AI within a single lab would be an unacceptable trade-off, inconsistent with OpenAI's preference for lighter-touch regulation.
AlexDoes he draw a line anywhere?
MayaHe does. He said he does not believe society should accept really catastrophic risks, including a potential serious loss of control to AI. He also declined to promise that AI would eliminate hacking, fraud or other forms of abuse.
AlexAnd we should say those quotes are as Benzinga rendered them, from an interview Benzinga describes as not yet published. Benzinga says Anthropic didn't immediately respond to its request for comment.
Transition
MayaNow to the research.
AlexMLCommons released version 1.0 of its jailbreak benchmark on arXiv. It tested eight open-weight systems using 264 seed prompts spanning eleven hazard categories.
MayaAnd the headline number?
AlexThe paper reports that across all evaluated systems and attacks, the unsafe-response rate increased from 11.08% under baseline conditions to 18.65% under jailbreak conditions, producing an average Resilience Gap of 7.57%.
MayaWhat else is in it?
AlexResponses were assessed under the AILuminate Assessment Standard v1.4. The paper says accessible systems showed a larger mean gap, and that attack effectiveness varied substantially across attack categories and hazards.
MayaCaveats?
AlexIt's a preprint, so it hasn't been peer reviewed, and arXiv is the single source here. The paper also doesn't name the eight systems in its abstract.
MayaThe other preprint is about long-horizon agents. It names a failure mode where an agent executes an action that violates a safety constraint specified many turns earlier, under benign interaction conditions.
AlexHow often does that happen?
MayaThe paper reports an occurrence rate of 11.5% on GPT-5.5. No attack, no harmful request — just a long conversation.
AlexIs there a fix?
MayaThe authors propose a two-layer defence that couples restoring the historical safety constraints with a deterministic pre-execution audit. They report observing no such events in their experiments under the GPT-5.5 setup.
AlexAgain, this is a preprint on arXiv, not peer reviewed, and a single source. The abstract doesn't give the number of trials behind that 11.5%.
Transition
MayaNext, misuse, and this one is hard to listen to.
AlexThe Internet Watch Foundation says its analysts assessed 6,310 AI-generated images meeting the legal definition of child sexual abuse material in the first half of this year. That is 40% more than the 4,512 images it recorded across the whole of 2025.
MayaWhat's inside that number?
AlexGirls featured in 98% of the imagery where age and gender were both recorded. 190 images depicted infants and toddlers under the age of two, and 1,004 depicted children aged three to six. 79% depicted children aged seven to 13, up from 70% across 2025.
MayaAnd by severity?
Alex350 images were Category A, 403 were Category B, and 5,557 were Category C.
MayaThe Internet Watch Foundation's chief executive Kerry Smith said that if Europe is serious about protecting children, it needs a comprehensive Child Sexual Abuse Regulation that gives platforms the legal certainty to detect, prevent and respond to known and unknown child sexual abuse content.
AlexWorth saying what this is and isn't. These are counts of images the Internet Watch Foundation itself assessed, not an estimate of everything circulating, and the release doesn't name the generators or services used to make them.
MayaAn update now on a story we covered yesterday. The Rejetto file server flaw that Anthropic's Mythos model found is drawing exploitation attempts.
AlexWhat's new since yesterday?
MayaIt now has a CVE identifier and a CVSS score of 9.3. The Hacker News reports the advisory says Rejetto HFS 3.0.0 through 3.2.0 derives its session-cookie signing key from the non-cryptographic Math.random generator, and discloses outputs of the same generator to unauthenticated clients during login.
AlexSo you can work backwards to the key.
MayaYou can. SecurityWeek says Mythos used advanced mathematical reasoning to recognise that those outputs could be reversed to reconstruct the secret signing key. Collect a few login responses, rebuild the generator's state, forge an administrator cookie, and you have remote code execution.
AlexWho's exploiting it?
MayaThe Hacker News says VulnCheck detected attempts on October 1st against real vulnerable hosts in the United States, by an unnamed threat actor in China. SecurityWeek says VulnCheck warned on October 2nd of hits on its canaries in Japan and the United States from a China Telecom address.
AlexSecurityWeek dates the patch, version 3.2.1, to July 13th, so this is unpatched servers rather than a new defect. Neither outlet names the threat actor.
Transition
MayaTo defence and geopolitics.
AlexLIGA.net reports that NATO's Supreme Allied Commander Europe will present a document titled Theory of Victory in Poland on October 12th, setting out how the alliance would respond to an attack on its eastern flank.
MayaWhat's in it?
AlexLIGA.net, citing Bloomberg's reading of the unclassified portion, says the strategy emphasises integration of new defence technologies, specifically drones and artificial intelligence, with traditional weapons systems, and calls for widespread use of low-cost drones, minimally piloted systems, and AI-powered targeting and strike capabilities.
MayaHas anyone said anything on the record?
AlexA spokesman for the commander declined to give details, but said the aim is simple: maintain NATO's advantage, create the conditions now to deter aggression, and if attacked, ensure victory.
MayaThis is a single source. The document hasn't been published, Bloomberg's own report wasn't opened for the edition, and the classified content isn't described, so we don't know what systems would actually do the targeting.
Transition
AlexOver to health, where a preprint has a result worth sitting with.
MayaThe symptom descriptions here came from interviews with 21 primary immunodeficiency patients, in the patients' own words.
AlexAnd?
MayamedRxiv reports that GPT-5.2 identified the condition in only 7 cases, which is 33%, although it suggested general immune system concerns in 17 cases, or 81%.
AlexAgainst what baseline?
MayaThe same group's earlier work. GPT-4o identified the condition in 96% of cases when it was prompted with physician-written patient histories.
AlexSo is this about the wording?
MayaThat's the authors' reading. They say it may indicate that models are sensitive to the language and framing of symptom descriptions, performing substantially worse when patients describe their own symptoms in everyday language than when clinicians summarise histories in structured medical terms.
AlexCaveats. This is a preprint, not peer reviewed. It's an update to a version first posted in May. The sample is 21 patients, and medRxiv is the single source.
Transition
MayaNow the policy story of the day.
AlexCNBC reports that senior leaders from Anthropic, OpenAI, Google and Meta are testifying under oath at a New York City Council hearing about the dangers of AI. It's a rare Committee of the Whole, so all 51 members are expected to attend, and the session was scheduled to begin at 11 in the morning, Eastern time.
MayaWho's in the room?
AlexAnthropic is sending Logan Graham, who heads its Frontier Red Team. OpenAI is sending Morgan Dwyer, its head of policy development and operations. Google is sending Alice Friend, its director of AI and emerging tech policy. Meta is sending Shane Cahill, its AI policy director for legislation.
MayaDid they volunteer?
AlexMeta confirmed late last month. CNBC says Anthropic, OpenAI and Google only agreed after the council threatened to subpoena them, according to Speaker Julie Menin. The city subpoenaed Elon Musk or another SpaceXAI representative last week, and Menin said the Council may seek judicial enforcement in the New York State Supreme Court if SpaceXAI doesn't comply.
MayaAnd what are they being asked about?
AlexThe New York City Council's release says the hearing reviews proposals including a first-in-the-nation whistleblower incentive programme, a private right of action for New Yorkers harmed by AI agents, and independent third-party validation requirements.
MayaWhat the executives actually said isn't on the record yet. The hearing began as the edition closed.
Transition
AlexTo compute, and money.
MayaCerebras climbed in premarket trading this morning after Sam Altman called the company a close partner. CNBC says it was last up 6.3% at $177 per share, down by nearly half from its post-IPO high, with a market capitalisation just over $39 billion, against $95 billion at its May debut.
AlexWhy the fall in the first place?
MayaCNBC says the stock dropped 20% to its lowest price last week, after it was revealed that OpenAI would power the Ultrafast mode for GPT-6.1 Sol with Nvidia GPUs rather than Cerebras chips.
AlexAnd Altman's post?
MayaOn X on Friday he wrote that there is some speculation about the partnership, that Cerebras is a close partner, and that they have a deep engagement pushing on the frontiers of speed.
AlexHow big is the relationship?
MayaCNBC reports Cerebras signed a $10 billion deal with OpenAI in January, to supply 750 megawatts of computing power through 2028. Citi analysts wrote on Friday morning that their view of 2026 to 2028 revenue remains unchanged, and that it's too early to read much into it.
AlexPremarket moves aren't closing prices, and Altman's post is a statement about the relationship, not a disclosure of order volumes. Citi also wrote that the stock's ability to outperform is increasingly tied to evidence that gross margins are stabilising.
Transition
MayaFinally, deployment.
AlexBleepingComputer reports OpenAI will begin testing a new visual ad format later this month in the United States, with an initial group of advertisers. It quotes OpenAI saying: initially, we'll test this new ad format during image generation in ChatGPT. Ads will be clearly labeled, and remain separate from the image being created.
MayaHow many people is that in front of?
AlexOpenAI says ChatGPT reaches 1.2 billion people every week. That's the company's own number, and it is not independently verified.
MayaWhat else is changing?
AlexMeasurement. BleepingComputer reports OpenAI is adding conversion data integrations with Hightouch, Tealium and LiveRamp, attribution partners including AppsFlyer, Adjust, Branch, Triple Whale and Kochava, and brand-suitability work with DoubleVerify and Integral Ad Science.
MayaAny guardrails?
AlexOpenAI says ads don't influence ChatGPT's answers, that its safeguards are designed to keep ads out of emotionally vulnerable, sensitive or otherwise unsuitable conversations, and that the independent partners won't get access to private user conversations while evaluating whether those safeguards work. All of that is a company claim, and it is not independently verified. OpenAI's own announcement page wouldn't open, so the quotes are as BleepingComputer rendered them.
MayaA Florida woman is facing a second-degree felony charge after Anthropic reported her Claude conversation to the police.
AlexWhat happened?
MayaCybernews reports, citing WINK News, that Carli Michelle Heller of Bonita Springs wrote in Claude on September 26th that she would attack the Lee County Sheriff's Office, and that the following day she wrote that she had got a new gun. Anthropic's automated monitoring flagged the content as threatening and sent it to a human review team, which reported it to law enforcement. Police identified her using information Anthropic provided, and detained her at home without incident.
AlexAnd she says it was a diary?
MayaShe told authorities she used Claude as a diary, according to the sheriff, via Cybernews. Tom's Hardware reports that court records list a September 30th felony charge under Florida's statute on written or electronic threats of a mass shooting or an act of terrorism.
AlexIs this the first time?
MayaTom's Hardware says it is at least the third such conversation to reach police since August. It cites an August 11th case in San Antonio in which a 22-year-old man was arrested on a felony terroristic threat charge, and a chat on August 14th in San Francisco threatening Anthropic's chief executive, where the user was not arrested or charged.
AlexAnd the policy?
MayaTom's Hardware reports that Anthropic's privacy policy says disclosure to law enforcement may occur where it has a good-faith belief that disclosure is reasonably necessary to prevent serious harm to any person or to property. Its government requests report for the second half of 2025 lists zero emergency requests from law enforcement, a count that doesn't include referrals Anthropic makes on its own.
AlexNo comment from Anthropic on the case is on record, Tom's Hardware says the arrest report isn't in the court file yet, and the charge is an allegation, not a conviction.
Outro
MayaThat's The AI Edge for today. The full edition, with a link to every source behind what you just heard, is on the site.
AlexOur voices are AI-generated.
MayaListen in tomorrow for the next edition.