← AI Switchboard
AI Switchboardby Waggle
Security · September 22, 2026
Sep 2

The labs shipped cyber-capable models — and their own guardrails

Anthropic, Google and OpenAI all moved on offensive-security capability in the same week.

September 2 was the day the frontier labs all moved on offensive security at once. Anthropic released Claude Fable 5.1 and Mythos 5.1 with new enterprise safeguards, Google put out Gemini 3.8 Flash Cyber for vetted defenders only, and OpenAI said its next model had reached the “Critical” cyber-risk level of its own Preparedness Framework — the first to do so.

Three weeks later the same capability class is the subject of a UN brief, a California executive order and two lab disclosures about test environments that got loose. The gap between shipping the capability and agreeing how to watch it is the whole story of this month.

  • Reported Anthropic released Claude Fable 5.1 and Mythos 5.1 with new enterprise safeguards; Google released Gemini 3.8 Flash Cyber for vetted defenders; OpenAI said its next model hit its “Critical” cyber-risk level. The Hacker News
Sources: The Hacker News · CNBC

Safety, security & governanceModels & releases

This month in the September 22, 2026 edition · front page