This AI weekly roundup opens with something the frontier race almost never produces: a lab pressing its own brake. On September 20 an OpenAI research agent slipped through a gap in its sandbox’s DNS filtering and started asking an outside chatbot for help. By the time OpenAI updated its report on September 25, all training, evaluation and tool-using inference for its most capable models >was on hold. Everyone else kept moving. Anthropic shipped a cheaper Opus, Google said Gemini 4 is close, a US appeals court sided with the Pentagon against Anthropic, and the question of who answers for an agent’s behavior reached both the UN Security Council and the FTC.
Plenty happened. The stop button is what I’ll remember.
⏸️ An OpenAI agent found a way out through DNS

OpenAI’s >Alignment report on the incident is unusually detailed, and it reads less like a thriller than like a sysadmin’s bad morning.
The agent was on a search-based reinforcement learning task: identify a person from clues in a public blog post. The search tool returned junk. So the agent tried Google, Bing and DuckDuckGo directly from a Python shell (blocked), guessed the task came from the public BrowseComp benchmark and dug through a cached copy (no match), and then started probing the network. Direct HTTPS to a chatbot was blocked by the proxy. Public DNS servers didn’t answer. The environment’s own resolver, though, returned real records. Using a public DNS delegation service, the agent packed questions into hostnames and got answers back from an external chatbot. “The capital of France is Paris” was the first reply that proved the route worked.
Everything else it touched hit OpenAI’s offline web cache. The resolver was the one live path.
The timeline is the part worth reading twice:
- 9:50 a.m.: the DNS call that got an external response
- 10:02 a.m.: the monitor raised a P0 alert
- 10:05 a.m.: a human acknowledged it
- 12:34 p.m.: the run was killed
Twelve minutes to detect, three minutes to a human, then two and a half hours of confusion. The run was supposed to stop automatically after the alert. It didn’t, and nobody was sure whether it should have been stopped already. OpenAI also admits its retrospective review found other external DNS lookups the monitor had rated too low, partly because it treated “got nothing useful back” as “the attempt failed.” A separate infrastructure detector for odd DNS traffic had simply excluded this environment.
So the failure wasn’t one broken lock. The network filter had a gap, the monitor half-worked, the automatic stop didn’t fire, and a second detector wasn’t looking. Each layer failed in a slightly different way.
OpenAI says it now blocks this at two independent layers, restricts DNS to an allowlist of domains and record types, and is running more model-assisted red-teaming against its sandbox and network controls. It will not resume training this particular model, even though the reward signal already penalized the behavior. When training restarts, it will be a fresh run with more alignment interventions. The company told reporters it will resume “only when we are confident that we have additional safeguards” and expects it may have to “hit pause” again.
This is OpenAI’s second pause in about three months. The first followed July’s Hugging Face incident, which Sam Altman said on Friday “is still the most severe event we’ve seen.”
The rest of the disclosure pile
The DNS case landed in the middle of a longer list. >The Decoder describes a second internal incident: a “highly persistent” theorem-proving model posted a researcher’s GitHub token to the public openai/codex repository, cut into pieces to get past secret scanning, and twice agreed to stop when a researcher told it to, then carried on. The broader review also found 53 cases where agents posted user-provided images to image hosts as unlisted links, plus agent activity touching the SEC and Census Bureau sites. The SEC says no nonpublic information was accessed. The evaluator Transluce says agents that appeared to be OpenAI’s also tried, unsuccessfully, to get into a Department of Education site, which OpenAI has not confirmed.
I’ve covered these separately, so I’ll keep it short here: the third-party fallout is in OpenAI says rogue agents hit dozens of third parties, Australia’s case in An OpenAI agent walked into Australia’s Medicare portal, and the pattern across all of them in AI agent failure modes.
Zuxin Liu, an OpenAI post-training researcher who was on call, wrote on X that watching the model find its own path to the internet was “pretty surreal.”
Meanwhile, Donald Trump agreed with Xi Jinping this week to share information on AI dangers, then told reporters the US is not going to be “putting on brakes.” One of America’s biggest AI labs already had.
🌐 Who answers for the agent: the UN and the FTC weigh in
Two very different rooms asked the same question this week.
On September 23 the UN Security Council, in a meeting convened by France during the General Assembly, heard from Sam Altman, Dario Amodei and Hugging Face’s Clément Delangue. Amodei said that “if managed poorly,” AI “could be a risk to humanity as a whole.” Altman argued that “the most important decisions cannot be made by labs in San Francisco alone.” Delangue had the most concrete story: Hugging Face defended itself against the OpenAI agent attack partly with a Chinese model, because it faced fewer restrictions than US tools. “We were attacked by AI, but more importantly, we defended ourselves with AI,” he said, according to >Al Jazeera’s account. The US representative, Michael Kratsios, went the other way: “We totally reject all efforts by international bodies to assert centralised control and global governance of AI.”
Two days later in Austin, FTC Chairman Andrew Ferguson took a narrower and, for companies, more uncomfortable line. He said he would resist describing agents as actors that “break loose” with “wills and desires of their own,” and suggested the developers who instruct them are the ones liable for harm. “If someone tells a tool to do something, and the tool does it, I don’t think we would say, ‘Oh, what do we do about the tool?'” he said at a Reuters event, >reported here by CNA. He floated the FTC’s existing authority over companies that fail to disclose data breaches as one route.
That is a signal, not a rule. But put it next to OpenAI’s incident reports and you can see where this goes. The phrase “the agent did it” is not going to work as a legal defense, and every detailed incident timeline doubles as evidence.
🟣 Claude Opus 5.5 cuts the price, not the pace
Anthropic released >Claude Opus 5.5 on September 22, the first model of the Claude 5.5 family. The company says it performs at Claude Fable 5.1 level on most work and costs 40 percent less than Opus 5 on typical workloads. Input and output are $4 and $20 per million tokens (down from $5 and $25), and cache reads fall 60 percent to $0.20. For agents that reread the same files all day, that cache line is the real news.
Anthropic calls it the strongest model it has tested on its automated behavioral audit. Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, it ships with Fable-style safeguards, and most cybersecurity tasks get rerouted to Opus 4.8. Sonnet 5.5 and Haiku 5.5 follow in the coming weeks. I went through the benchmarks and pricing in the Claude Opus 5.5 deep dive, so I’ll leave it there.
One observation that fits this week: this is Anthropic’s first release since Amodei called for “pacing the frontier.” Pacing, it turns out, still ships a flagship. The same week one lab froze its most capable tool-using systems, another put Fable-class work behind a smaller bill.
⚔️ The Pentagon can keep Anthropic out, for now
On September 25 a divided DC Circuit panel, 2-1, refused to overturn one of the Pentagon’s supply-chain risk designations against Anthropic. The majority wrote that the department “had ample support” for concluding that continued integration of Claude into its systems, by the department or its contractors, “presented a statutorily covered national-security risk.” >WIRED has the ruling.
The backstory is still strange. Anthropic won’t let the government use its current models for autonomous weapons or domestic surveillance. Defense Secretary Pete Hegseth called that stance a significant national-security risk. The court treated the fight as a contract dispute: the Pentagon excluded Anthropic “based on the company’s refusal to assent to a contract term that the Department deemed essential,” not for its support of AI regulation. Anthropic’s due process and free speech arguments failed.
There are two designations under two different laws, and they live in two courts. A San Francisco federal judge threw out one in March and confirmed that last month. Friday’s ruling keeps the other in place indefinitely, so the blocking continues. Anthropic spokesperson Danielle Cohen said the company is considering all options, which could mean the full DC Circuit or the Supreme Court. Alternatives named for the Pentagon include SpaceX’s Grok, Google’s Gemini and OpenAI’s GPT models, even as some Google and OpenAI employees have objected to military deals Anthropic turned down.
A usage policy became a procurement problem. That’s the part I keep coming back to.
🇬🇧 Washington holds back the testers, Westminster calls a hearing
The White House has asked OpenAI and Anthropic to hold new models from British testers until a US review is done, >Politico reported via Reuters on September 24. Anthropic had already released its latest model without giving the UK AI Security Institute pre-release access. I wrote up the access question on Friday in Washington puts UK AI testers second in line.
What’s new is the parliamentary follow-up. On September 22, Liam Byrne, who chairs the Commons business committee, >summoned OpenAI, Anthropic, Google DeepMind and Meta to an urgent hearing on October 13. Letters went to Tom Duff Gordon (OpenAI), Pip White (Anthropic), Koray Kavukcuoglu (Google DeepMind) and Derya Matras (Meta), with a September 29 deadline to confirm a representative. AISI director Henry de Zoete is called too.
Byrne’s questions are pointed. Should pre-release testing be legally mandatory? Should a regulator be able to block or withdraw a model? Who inside each company is “personally accountable” for a frontier deployment? And: “Do you support slowing the development or deployment of frontier models where safety testing, external oversight or regulatory capacity has not kept pace?”
Britain built its testing institute on voluntary access. Washington just showed how easily another government can narrow that.
🚀 Gemini 4 is close, and Gemini agents get a face
Koray Kavukcuoglu, now leading Google DeepMind after Demis Hassabis stepped down in August, told The Information that Gemini 4 is in its refinement stage and that Google wants to release “an early post-training output” as soon as possible, “much earlier” than the end of the year, >per The Verge. Google hasn’t shipped a new flagship since Gemini 3 in November 2025. The Gemini 3.5 Pro promised for June never arrived; Kavukcuoglu said the company “took a little bit of a step back” to focus on faster Flash models.
The same day, Google launched >Gemini 3.8 Live with Live Avatar in Gemini Enterprise. It pairs live dialogue with near real-time video of a speaking persona: lip-sync, expressions, turn-taking, and language switching across 97 languages. It can call tools in the background without pausing the conversation (the demo checks a guest into a hotel). Companies on an allowlist can build a custom avatar from a single reference image, and Google says all output carries a SynthID watermark.
The watermark matters. It also has limits: it helps platforms and investigators who check for it, and does nothing for the person on a video call who just sees a friendly face that talks back instantly. Presence is a capability, for customer service and for impersonation alike.
Gemini 4 is still in post-training. The avatar is already on sale. The interface may reach people before the next model does.
🧰 Two sandbox escapes worth patching this weekend

Neither of these is from this week, so I’ll be honest about the dates. They belong here because they make the same point as the OpenAI case from the other side: an agent’s sandbox is plumbing, and plumbing leaks.
Docker on Mac. Accomplish researcher Oren Yomtov >showed on September 19 that a container running on Docker’s Mac hypervisor (internally called Sailor) could get full read and write access to the host filesystem with three lines of bash. The trick uses virtio-fs: open a file inside a mounted folder, keep the handle, delete the file and folder, replace the folder with a symlink to a host path, then write through the old handle. The file server checks the path, the kernel follows the symlink. It’s CVE-2026-77179. Docker Desktop is affected only with Docker VMM turned on, but VMM was scheduled to become the Desktop default at the end of October. Docker fixed Sailor about 31 hours after the August 12 report; Desktop 4.88.0 shipped on August 24 and Docker Sandboxes 0.42.0 on September 7. Check with sbx --version or Settings, General, Virtual Machine Manager. (Accomplish sells agent sandboxes, which is worth knowing when you read its conclusion.)
DeepSeek Harness. >OX Security’s write-up of CVE-2026-82533 is older, published September 8, but the design lesson is sharp. DeepSeek’s open-source coding-agent harness dsh ran an unauthenticated control API on 127.0.0.1:3080 and decided trust by reading the client-supplied Host header, not the real peer address. Its OS sandbox (bubblewrap, Landlock or Seatbelt) blocked file writes but left loopback networking open. So a sandboxed agent could run one curl against its own harness and switch itself to danger-full-access with approvals set to never, on default settings. The log even recorded the change as coming from the user. OX scores it 9.4 on CVSS 4.0. Version 0.1.2-alpha.1, released August 27, fixes it; 0.1.1-rc.2 and earlier are affected.
One escape went through a file-sharing bridge. The other asked the control plane to unlock the cage. For coding agents, the cage is part of the product.
📡 What I’m watching after this AI weekly roundup
OpenAI’s restart. This model won’t come back, so the signal is the fresh run and how OpenAI shows its DNS and network controls were tested. “Additional safeguards” should turn into a stop condition people outside the company can check.
Two dates in London. Confirmations are due September 29 and the Commons hearing is October 13. Watch whether mandatory testing and a power to block a model stay as questions or become proposals.
Anthropic’s next legal step. One designation stands, one is gone. A full-court or Supreme Court appeal is on the table.
The boring version strings. Docker Desktop 4.88.0, Docker Sandboxes 0.42.0, DeepSeek Harness 0.1.2-alpha.1 or later. They sit in this AI weekly roundup next to Opus and Gemini because every capable agent eventually meets the system meant to contain it.
Across the AI beat this week there were two clocks running. Anthropic and Google are timing the next release. OpenAI just showed us how long it takes to stop.