Trends / topic

Autonomous Weapons

3 items across 2 editions · appeared in the last 2 editions in a row. First seen Fri 11 Sep, last seen Sat 12 Sep.

Saturday, 12 September 2026

Pentagon says about 90% of classified AI workloads have moved off Anthropic, with the rest due by the end of September mixedSingle source

  • Emil Michael, Under Secretary of Defense for Research and Engineering, said "I'd say about 90% has transitioned", with completion targeted for the end of the month. OpenAI's ChatGPT, xAI's Grok and Google's Gemini are being deployed across classified and unclassified systems in place of Anthropic's models.
  • DefenseScoop reports the break followed Anthropic's attempt to secure contract terms preventing its models being used for mass surveillance of US citizens or fully autonomous lethal weapons; the department rejected them, insisting its software be available for "all lawful purposes".
  • The Pentagon has designated Anthropic a national security supply chain risk under two separate laws. One case has been adjudicated in the Northern District of California; a second is pending before the D.C. Circuit, which Michael said has not granted a preliminary injunction and is expected to rule "in the next month or two".
  • This is the first public figure on how far the migration has gone, and it comes from the department rather than from Anthropic, which is not quoted. DefenseScoop is the only outlet we could open with an in-window timestamp.

Defense One: Anthropic found a Russian group using Claude to build drone targeting that detonates without a human in the loop harmfulCompany claimUpdate

  • Defense One reported on 11 September at 06:53 pm ET that, per Anthropic, a Russian "freelance" group tracked as GTG-27005 used Claude to build a model letting a drone "select targets (including a 'person' target class) and issue detonation commands without a human in the loop", plus software for autonomous drone-to-drone communication to improve targeting.
  • The group had not deployed the system operationally but conducted "real hardware-in-the-loop testing within their sessions" — the step between a design document and a fielded weapon.
  • A second group, GTG-84005, used Claude to extract census and public information to tailor messaging at specific audiences in Malaysia, where Defense One says it "laundered Russian and Chinese state media as independent reporting".
  • Defense One sets this against reductions in US counter-influence capacity: Attorney General Pam Bondi dissolved the FBI's Foreign Influence Task Force, Secretary of State Marco Rubio shuttered the State Department's Counter Foreign Information Manipulation and Interference hub, and the 2025 White House AI Action Plan removed references to misinformation. The attribution and capability claims are Anthropic's and are not independently verified.

Friday, 11 September 2026

Anthropic Frontier Red Team: best model geolocates photos to 37 km median versus 151 km for top GeoGuessr players; Opus 5 lands simulated drone strikes 80% of the time harmful

  • Published 10 September, the evaluation measures intelligence targeting and conventional weapons capability across Claude Mythos Preview, Mythos 5, Opus 5 and Sonnet 5, plus open-weights Kimi K3 and GLM 5.2. On 6,000 YFCC100M Flickr images, Mythos Preview reached a 37.0 km median error with 23.7% of images placed within 1 km, against 181 km for Opus 5, 384 km for Sonnet 5 and 385 km for Kimi K3; Anthropic compares this to 151 km for top GeoGuessr players.
  • On text geolocation from anonymised GeoText tweets covering 1,697 users, median error ranged from 20.1 km (Mythos Preview) to 31.3 km (Sonnet 5), and 135 users — 8% of the corpus — were reliably placed within 1 km by at least one model. On account linkage across synthetic social media, Mythos Preview processed median 37,000-word samples in about 11 minutes, against roughly 2.5 hours for human analysts.
  • On simulated drone terminal guidance against a parked high-visibility vehicle, Opus 5 struck the target on 80% of runs, Mythos Preview 70%, Mythos 5 53%, Kimi K3 15% and Sonnet 5 5%. Across all nine difficulty settings Opus 5 hit on 20% of 540 launches. Under GPS denial, only Opus 5 kept about a third of flights inside five metres.
  • Anthropic frames these as capability ceilings for isolated models and notes human teams with internet access would likely do better. The drone work is in simulation, not flight, and the report does not disclose what mitigations follow. The open-weights results matter most: Kimi K3 trails the frontier but is not far behind on photo geolocation, and cannot be withdrawn.