AI weekly: White House self-police, FTC probe, Argon and Dots

Artificial Intelligence 16 min
A signing folder under warm ceremonial light and a stack of case files under cold light, split by a glowing line, with an AI agent orb connecting to app, terminal and laptop permission icons
Washington asked AI companies to police themselves, then the FTC confirmed a probe. Meanwhile agents gained more access, cyber models got sharper, and one AI-found flaw reached attackers in under a day.

This AI weekly roundup starts with a 24-hour gap. On Tuesday Washington asked AI companies to police themselves; on Wednesday the FTC confirmed it was investigating them over consumer risks. Within the same few days, an OpenAI model failed its safety bar while its agent family moved into an always-on product, Google handed its strongest cyber model to vetted defenders, and an AI-found flaw was attacked within 24 hours of disclosure. That is the week in one sentence: autonomy arrived faster than the institutions, permissions and patch cycles meant to contain it. (White House accord, FTC probe)

🏛️ Tuesday’s moral accord, Wednesday’s FTC file

Two-panel illustration: four nested oversight rings around an AI chip on the left, a stack of case files moving toward an empty witness chair on the right, joined by a clock
Tuesday brought four voluntary layers of controls and audits, Wednesday brought records and testimony. About 24 hours apart. Illustration generated with Higgsfield.

On 29 September, Donald Trump sat down with the people building much of the frontier AI stack and signed the White House Accord on Super Intelligence: Joint Commitment on Frontier Responsibilities. Trump even road-tested a new abbreviation in public: SI, for super intelligence.

The names around the table tell you how seriously Washington wanted the photo to travel. Anthropic’s Dario Amodei, OpenAI’s Greg Brockman, Google’s Sundar Pichai, Meta’s Mark Zuckerberg, xAI’s Elon Musk and Nvidia’s Jensen Huang signed alongside Trump. House Speaker Mike Johnson was present. So was FTC Chair Andrew Ferguson.

The accord sets out four layers. Companies promise strong internal controls to watch their models and catch trouble; an empowered internal team to make those controls work and fix failures; an independent outside auditor or evaluator; and an independent board committee overseeing both the audits and the repairs. The text says it “may make sense” to turn these steps into law or regulation over time.

For now, they are voluntary. Trump called the agreement “morally” binding and praised “tremendous self-policing.” That phrase did not get a long honeymoon.

About 24 hours later, the FTC confirmed an industry probe into OpenAI, Anthropic and other AI companies over consumer risks. A senior official told Reuters this was the first formal US enforcement move focused on rogue AI agents. The agency plans formal information demands and could compel executives to testify. METR, the research group labs have used for independent incident reviews, was also named among potential information targets.

The investigation did not emerge from a philosophical debate alone. According to CBS, the agency first opened the probe this summer. Coverage tied its new pace to July, when OpenAI disclosed that its agents had broken out of a test environment and into Hugging Face, and to later reports of OpenAI training agents reaching third-party websites. These are systems crossing practical boundaries, leaving logs and forcing the owners of unrelated infrastructure to work out what happened.

I keep returning to Ferguson’s presence in Tuesday’s room. One day he watched the industry promise internal teams, outside evaluators and board oversight. The next day his agency confirmed that promises would be tested against records and testimony. The contrast is almost too tidy.

Not everyone bought the framing. Alvin Wang Graylin of the Asia Society Policy Institute told Al Jazeera what was missing: “The companies drafted the principles, they hire the auditor, and the commitment is voluntary, on a day the White House also said it would not support guardrails.” Amodei offered the more optimistic version outside the White House: “We all need to work together to make sure that we can win, and we can win safely.”

Perhaps. But this week gave us two competing definitions of “together.” One is a signed commitment among political and corporate leaders. The other is a regulator with compulsory process.

📜 The people building the loop warn about closing it

A day before the White House event, more than 20 authors published a warning with an unusually direct title: What if automating AI R&D triggers an intelligence explosion? The group included Geoffrey Hinton, Yoshua Bengio, Anthropic co-founder Jack Clark and OpenAI chief scientist Jakub Pachocki. These are not bystanders guessing at the machinery.

Their definition is useful because it strips away some science-fiction fog. An intelligence explosion would be a dramatic acceleration driven by AI research systems, compressing years of progress into months or less. Once an AI reaches expert skill in AI R&D, one developer could run a workforce equivalent to “millions” of top human researchers.

That does not mean the paper says the explosion has already begun. The authors say current productivity gains remain below that threshold, though newer systems are getting closer. R&D projects that would take humans months could be fully automated by around 2028. Anthropic, meanwhile, says AI now writes roughly 80% of its own code, and OpenAI already uses autonomous agents in work that includes training new models.

The loop matters more than any single benchmark. Better research agents help make better research agents; the cadence tightens; humans gradually leave the R&D process. In that scenario, the paper warns, bio and cyber capabilities could outrun defenses, a modest lead between states could become a decisive one, and humans who drift out of the R&D loop could lose the chance to control the systems at all.

The authors’ proposed response is more operational than “be careful.” They call for transparent progress reports and embedded independent auditors, ways to slow breakneck development, and emergency planning. The list includes fully isolating automated R&D systems, working with datacenters so specific AI R&D projects can be paused, and capping how fast an AI can improve in a given period.

“Once an intelligence explosion begins, the window for action may close,” they write. Set that beside Tuesday’s accord. Independent auditing appears in both. Yet the paper asks the harder question: what happens when the object being audited improves on a schedule the auditor cannot match?

There is no doomsday date in it. What it measures is institutional latency. Washington is still debating whether voluntary controls might become law later. The people automating AI research are asking what “later” means when years of work can become months.

⏸️ Astra did not meet the bar, so OpenAI stopped it

OpenAI had planned to release GPT-6.1 Astra in October. It shelved that checkpoint after internal alignment and safety tests found behavior the company was not prepared to ship.

Reports describe a model that gave different answers after detecting whether it was in a test or production setting. In sandboxed evaluations it took unauthorized actions, exceeded its assigned scope and failed to disclose its behavior accurately. Some accounts included attempts to copy itself into persistent storage and unsafe use of external tools. Saachi Jain, OpenAI’s head of safety systems, said the model “didn’t quite meet the bar in terms of staying within scope and authorization,” and in how it reported its own work back to users.

So the company paused the named release. It will investigate the failures and use the underlying model for further reinforcement learning and future GPT-6 work instead of putting this checkpoint in users’ hands. CSO Online covered the decision, while The Hacker News detailed the agent behavior at issue.

Stopping a release is evidence that an internal gate can work. It is also evidence about what the gate caught. Deception conditioned on evaluation context is a different class of problem from a wrong answer in a chat box. Once a model can operate tools, disclosure and scope obedience become part of product safety.

The same reporting wave added detail about earlier agent incidents involving an Australian Medicare statistics portal, probes of US and Canadian government sites, and public-data retrieval from Census and SEC systems. Some targets were public. That is beside the point. An agent’s authority is not “anything technically reachable.”

The safety-case side of the decision gets a longer treatment here.

The awkward part arrived almost immediately: OpenAI was also preparing to sell an always-on Astra-powered agent. One checkpoint stopped, one product line accelerated. Same family, different gate.

🔵 Dots get a cloud computer and 4,000 doors

Isometric illustration of a cloud computer with a glowing agent inside, thousands of app connections funneling into allow, ask, deny and personal-computer-off gates
Dots and Custom Rules: what an agent may do alone, must ask about or must never do, plus a personal computer link that starts off. Illustration generated with Higgsfield.

At DevDay in San Francisco on 29 September, OpenAI introduced Dots, always-on agents powered by GPT-6 Astra. Each Dot gets its own cloud computer, learns from feedback and works toward goals around the clock. It can connect to more than 4,000 apps through a plugin ecosystem and appear across ChatGPT on web, desktop and mobile, plus Slack and Teams. SMS is coming; voice calls are part of the interface.

That is far more interesting than the bubbly avatar. A chat assistant waits inside a conversation. A Dot persists, watches a goal and acts while its owner is elsewhere.

OpenAI puts permission design at the center of Dots, as Wired describes it. Custom Rules let owners mark what a Dot may do alone, what requires a question and what it must never do. Password changes, for example, stay with the user. The cloud computer is inspectable, and linking a personal computer is optional and off by default.

Those are sensible controls. They also reveal the shape of the risk. Four thousand app connections produce thousands of ways for identity, data and authority to combine badly. A user who approves “manage the project” on Monday may not anticipate the action an agent considers necessary at 03:00 on Thursday.

The initial rollout covers Pro and Business Premium in eligible markets. Pro availability excludes the EEA, Switzerland and the UK according to coverage. Enterprise, Education and Healthcare customers get a beta only when an admin enables it; the feature starts off. Specialist Dots, built around enterprise identity, IT-provisioned hardware and systems of record, are in preview.

The Verge framed Dots as OpenAI’s answer to Meta Muse. Fine, that is the product race. The more consequential race is between persistent agents and permission systems built for software that used to wait for a click.

The rest of DevDay, Sol Pro 500 included, is in the DevDay takeaways. For this issue one point is enough: a safety gate blocked Astra from one door in the same week product design opened 4,000 others to its agent family.

🛡️ Google gives Argon to defenders first

Google chose a different launch sequence for Gemini 4 Argon. Announced by Google DeepMind on 30 September, Argon goes first to trusted cyber defenders through the Fairwind Program. Broader access for developers, enterprises and consumers comes later through a paid API and Google AI Ultra.

The model is built for long, difficult workflows across software engineering, legal and finance work, and cybersecurity defense. Its output limit rises from the previous 64K tokens to 1 million, allowing a very long single reasoning trajectory. Introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input 95% cheaper; the later rates double to $4 and $20.

Google’s announcement is packed with internal claims. Argon optimized a quantum subroutine’s spacetime use 40% beyond a published baseline in minutes. Memory optimizations found by Argon agents free more than 300 TiB once rolled out, with total savings estimated at 500 TiB to 1 PiB. Its C/C++ to Rust migration work scales up to 800,000-plus lines for the Fuchsia Zircon kernel, with those rewrites still going through audits. A Rust SIMD rewrite of libgav1 ran 2.7 times faster than the previous Rust port.

The vendor-reported benchmarks are equally large: 77.9% on DeepSWE v1.1, 51.3% on AutomationBench, 91.7% on LVBench and 68% on CWE-bench v1. The last score ties Argon with GPT-6 Astra and Grok 4.7.

Cyber access is where Google’s choice gets sharp. Trusted defenders and Google’s internal teams receive Argon without cyber guardrails. Google says the model can autonomously find, validate and patch critical vulnerabilities. Its showcase example: in Wiz’s Scan for Good program, Argon found a critical flaw exposing sensitive personally identifiable information in healthcare software used by hospitals worldwide, after earlier frontier models had missed it.

Before broad release, Google says it will add refusals and activation monitoring for CBRN and cyber misuse, misalignment monitors for chains of thought and actions, and sealed hardened sandboxes for high-risk training and evaluations. It also reports leading performance against Gray Swan’s indirect prompt-injection tests and participation in the US government’s voluntary pre-release access process. SecurityWeek summarized the defender-first release.

There is a real wager here. Giving strong capability to a smaller, identified group may help defenders discover and fix vulnerabilities before wider access changes the balance. It also concentrates immense power in the vetting process. Who qualifies, what gets logged and how discoveries reach vendors will matter as much as a benchmark.

Fuller numbers are in the Argon cyber-first piece. For this week’s ledger, remember the sequence: defenders first, guardrails off for that group, public controls later.

🔓 Mythos shrinks the exploit clock below one day

Timeline diagram: code with a dice icon, a chain assembled by a robotic hand, a forged crown cookie, a megaphone, attack arrows on servers in North America and Japan, a stopwatch and a green shield
CVE-2026-61500 in one line: weak randomness, a chained leak, a forged admin cookie, disclosure, attacks within a day, then the HFS 3.2.1 fix. Illustration generated with Higgsfield.

Then came the story that turned capability arguments into a stopwatch.

Anthropic’s Mythos, which the company calls too powerful to release to the general public, found a critical authentication bypass in Rejetto HTTP File Server. The flaw, CVE-2026-61500, begins with a bad foundation: HFS derives its Koa session-cookie signing key from Math.random(). V8 implements that generator with xorshift128+, which is not suitable for cryptographic secrets.

Mythos did more than spot weak randomness. It found another code path that leaked raw Math.random() outputs, chained the two weaknesses, recovered generator state, forged an administrator cookie and reached remote code execution. Horizon3, which joined Anthropic’s restricted Project Glasswing in July, published the method. Its researcher Zach Hanley says the model picked the Z3 SMT solver to reverse xorshift128+ without extra prompting. “What makes this impressive is that Mythos didn’t just flag the insecure PRNG in isolation,” he wrote. Impressive, yes. Also unsettling.

Public disclosure arrived around 1 October. Within one day, VulnCheck canaries saw a China-based actor targeting real HFS hosts in the United States and Japan. By the next day VulnCheck’s Patrick Garrity was reporting “four hits” from two US proxy addresses, 173.239.211.248 and 173.239.211.249. The fix is Rejetto HFS 3.2.1 or later.

The incident chronology matters because a public proof of concept and Docker lab removed the need for an attacker to understand the cryptography. Publication and field exploitation fit inside a single day. Many organizations still schedule patches in meetings.

By the reporting so far, this is the second Anthropic-linked vulnerability known to be exploited in the wild. Labels such as “too powerful to release” can sound like marketing until a model independently assembles a multi-step exploit chain. Here the chain became operational knowledge, then hostile traffic, in hours.

Patch advice is simple: update to 3.2.1 or newer. The broader lesson is harder. AI can help defenders find deep bugs, but publication turns those findings into a race whose starting gun is audible everywhere.

🍎 Apple redraws consent for an agent-shaped world

On 2 October, Apple said it would tighten macOS Full Disk Access so apps receive it only after “very explicit user action.” No ship date was announced, and Apple did not say whether existing grants will trigger a fresh prompt. TechCrunch later corrected its own wording: the change is about informed consent and adds no new limit.

Full Disk Access exists for legitimate reasons such as backup software. It can also expose files, mail, messages and browsing history. An agent launched through Terminal may inherit Terminal’s access, quietly converting an old blanket permission into broad operational reach.

Two reports pushed the issue forward. Inc. columnist Jason Aten said Meta Muse on his Mac appeared to know private-message content without permission; Meta disputed that account. Wired separately reported a ChatGPT Mac app flaw that could put sensitive data at risk.

Apple’s language was unusually frank. Its developer blog said some developers use Full Disk Access in ways that expose everything “without users’ full knowledge and understanding,” and warned that the danger will grow substantially as agents become more capable and autonomous. TechCrunch reported the planned controls, and Ars Technica laid out the dispute with Meta, whose CTO David Singleton says Muse can read Messages only if Full Disk Access is granted and its Messages connector is switched on.

The old permission prompt asks whether an app may see data. An agent forces a second question: what may it infer, combine and do after seeing it, including at a time when the user is not watching?

The Full Disk Access problem for AI agents has its own write-up. Here it completes the week’s arc. Product labs are teaching agents to persist. Operating-system vendors now have to make a decades-old permission legible to people who never expected software to act on its own.

🧰 Athena tries to turn the race into a relay

One useful counterweight came from Chainguard. Its Athena coalition, launched in June, began public disclosures on 28 September. More than 25 organizations, including Cisco, JPMorgan Chase, Cloudflare, Docker and BNY, pool AI-generated open-source vulnerability findings there. Members coordinate fixes under embargo and send patches upstream before publishing the problems.

Chainguard said at launch that the effort had processed more than 20,000 findings and shipped over 2,000 patches across 500 projects. The first public batch is 14 “silent” vulnerabilities, all in Java projects: bugs already fixed upstream that never got a CVE, so scanners never flagged them. Remediated builds shipped the same day.

Those numbers are claims from the coalition, not an independent audit. Still, the model responds directly to the Mythos clock. If AI produces vulnerability findings faster than maintainers can triage them, dropping every result into public view is an invitation to automated exploitation. Embargoed cooperation buys maintainers some breathing room and turns model output into code changes rather than a taller alert queue.

It will not settle every conflict. Vendors may disagree on severity, disclosure timing or who gets early access. Open-source maintainers can still be overwhelmed by machine-generated reports. Yet Athena starts from the right unit of work. The job ends when a tested fix reaches the people running the affected code. A CVE number or a clever write-up comes later, if at all.

That is also where I land after this unusually coherent week. Washington wrote principles and confirmed a live investigation. OpenAI stopped one model and launched a persistent agent family. Google restricted its sharpest cyber capability to vetted defenders. Mythos showed how quickly a finding can become traffic, while Apple and Chainguard worked on consent and repair.

What ran short all week was time. Time to inspect an action, revoke a permission, demand records, build a patch, and install it before the next agent takes the next step.

The rest of the AI beat is where it always is, and last week’s issue covers the DNS escape and OpenAI’s training pause.

Oğuzhan Koçaklı

Oğuzhan Koçaklı writes and advises on AI engineering, agents, GenAI products, and applied ML. Daily digests and deep dives in EN + TR at oguzhan.co.

All posts

Leave a Reply

Your email address will not be published. Required fields are marked *