The GenAI Field Guide · Trending
The three largest labs now ship their strongest models in two versions: one with cyber safeguards for everyone, and one with those safeguards relaxed for vetted defenders only.
Between February and September 2026, Anthropic, OpenAI and Google each settled on the same release pattern for models that can find and exploit software vulnerabilities: a general release with cyber classifiers, plus an application-only programme (Project Glasswing, Daybreak, Fairwind) where verified security teams get the same or a purpose-trained model with fewer refusals. Google's Gemini 4 Argon on 30 September went further and reached Fairwind defenders before any public release. If you build on these models, expect cyber refusals in ordinary engineering work, and if you do security work, the strongest capability now sits behind an identity check rather than a price list.
This is a pattern, not a product. We are grouping it as one trend; each lab describes its own programme and none of them calls it an industry approach. The common shape is that capability is no longer the same for every customer of a model. The question that decides what the model will do for you has shifted from which model you call to who you are verified to be.
Anthropic's version: Claude Fable 5.1, released 1 September 2026, is generally available with cyber safeguards. Claude Mythos 5.1 is, per Anthropic's docs, the same model with the same specifications and price ($10 per million input tokens, $50 per million output), offered by invitation only to Project Glasswing participants. Anthropic's announcement says Mythos 5.1 has reduced cyber safeguards for defensive security work and is limited to vetted cyber defenders and life scientists in US organisations. Glasswing itself dates from 7 April 2026, when Anthropic said it did not plan to make Claude Mythos Preview generally available and named launch partners including AWS, Apple, Cisco, CrowdStrike, Google, Microsoft and the Linux Foundation.
OpenAI's version is a programme called Trusted Access for Cyber, now branded Daybreak, with two tiers. Per OpenAI's own documentation, Daybreak Blue gives approved defenders flagship models with fewer refusals for defensive work such as vulnerability discovery and incident response, and Daybreak Red is a separately approved tier for specialist cyber models used in exploit validation, penetration testing and red teaming. OpenAI's deployment safety page for GPT-6 Astra (3 September 2026) says Astra meets its Critical cyber threshold, the first OpenAI model to do so, and that cybersecurity access is restricted to qualified professionals.
Google's version is the Fairwind Program. Google's 30 September post says Gemini 4 Argon is rolling out first to trusted cyber defenders through Fairwind, that Google is taking part in the US government's voluntary pre-release access process, and that for trusted defenders and its own internal teams it will release Argon without cyber guardrails. Developers, enterprises, paid API customers and Google AI Ultra subscribers are promised access as soon as possible, with no date given.
Each gate has two parts: a safeguard layer on the general release, and a verification process that relaxes it. On the Claude API, the safeguard layer is visible. Anthropic's refusals documentation says Fable 5.1, Fable 5, Opus 5.5, Opus 5 and Sonnet 5.5 carry safety classifiers that return an HTTP 200 response with stop_reason set to refusal and a stop_details.category such as cyber. The docs state plainly that benign cybersecurity work can also trigger the cyber category. A cyber refusal before any output is not billed, and an optional server-side fallback (beta) can retry a refused request on a model Anthropic recommends for that category, which may be less capable.
Anthropic's announcement says the Fable 5.1 classifiers block penetration testing, exploit generation, binary-based vulnerability scanning and malicious agentic coding, while allowing defensive identification of vulnerabilities, and that Claude Code users should see around 60 percent fewer cyber interventions per session than under Fable 5's safeguards. Its Cyber Verification Program gives verified users certain Opus and Sonnet class models with reduced cyber safeguards, with Mythos class models described as coming soon.
OpenAI's verification is the heaviest of the three as described. Its documentation says Daybreak access is controlled through identity verification, account security, monitoring, approved-use restrictions and legal attestations, that individuals apply through ChatGPT and organisations through a form, and that applying does not guarantee approval. Forkast reports that since 1 September 2026 every Daybreak Red account must also have a hardware security key; we could not confirm that on an OpenAI page. On the safeguard side, OpenAI's GPT-6 Astra card describes monitoring of complete trajectories including chains of thought, with flagged items scored for human review, and OpenAI's docs say a safeguard can block, reroute or limit a request.
Google's Fairwind page asks partners to restrict access to internal security, incident response or penetration testing teams, to use user-level authentication with phishing-resistant MFA, and to use the models only for authorised defensive and academic work. Priority goes to governments and national cyber authorities, critical infrastructure operators and core technology platforms. Google says Argon's general release will refuse harmful requests while preserving dual-use research under its Frontier Safety Framework, with monitoring of internal activations to spot misuse and of the model's chain of thought and actions for misalignment.
Most readers meet this trend from the general-release side: a request about a CVE, a fuzzing harness or a log full of exploit strings comes back refused. You need nothing new to handle that. To get the relaxed tier you need to be doing authorised security work, usually inside an organisation, and to pass a review that none of the three labs promises to complete on a timeline.