The GenAI Field Guide · Trending

Gated frontier releases for cyber capability

The three largest labs now ship their strongest models in two versions: one with cyber safeguards for everyone, and one with those safeguards relaxed for vetted defenders only.

Between February and September 2026, Anthropic, OpenAI and Google each settled on the same release pattern for models that can find and exploit software vulnerabilities: a general release with cyber classifiers, plus an application-only programme (Project Glasswing, Daybreak, Fairwind) where verified security teams get the same or a purpose-trained model with fewer refusals. Google's Gemini 4 Argon on 30 September went further and reached Fairwind defenders before any public release. If you build on these models, expect cyber refusals in ordinary engineering work, and if you do security work, the strongest capability now sits behind an identity check rather than a price list.

What it is

This is a pattern, not a product. We are grouping it as one trend; each lab describes its own programme and none of them calls it an industry approach. The common shape is that capability is no longer the same for every customer of a model. The question that decides what the model will do for you has shifted from which model you call to who you are verified to be.

Anthropic's version: Claude Fable 5.1, released 1 September 2026, is generally available with cyber safeguards. Claude Mythos 5.1 is, per Anthropic's docs, the same model with the same specifications and price ($10 per million input tokens, $50 per million output), offered by invitation only to Project Glasswing participants. Anthropic's announcement says Mythos 5.1 has reduced cyber safeguards for defensive security work and is limited to vetted cyber defenders and life scientists in US organisations. Glasswing itself dates from 7 April 2026, when Anthropic said it did not plan to make Claude Mythos Preview generally available and named launch partners including AWS, Apple, Cisco, CrowdStrike, Google, Microsoft and the Linux Foundation.

OpenAI's version is a programme called Trusted Access for Cyber, now branded Daybreak, with two tiers. Per OpenAI's own documentation, Daybreak Blue gives approved defenders flagship models with fewer refusals for defensive work such as vulnerability discovery and incident response, and Daybreak Red is a separately approved tier for specialist cyber models used in exploit validation, penetration testing and red teaming. OpenAI's deployment safety page for GPT-6 Astra (3 September 2026) says Astra meets its Critical cyber threshold, the first OpenAI model to do so, and that cybersecurity access is restricted to qualified professionals.

Google's version is the Fairwind Program. Google's 30 September post says Gemini 4 Argon is rolling out first to trusted cyber defenders through Fairwind, that Google is taking part in the US government's voluntary pre-release access process, and that for trusted defenders and its own internal teams it will release Argon without cyber guardrails. Developers, enterprises, paid API customers and Google AI Ultra subscribers are promised access as soon as possible, with no date given.

How it works

Each gate has two parts: a safeguard layer on the general release, and a verification process that relaxes it. On the Claude API, the safeguard layer is visible. Anthropic's refusals documentation says Fable 5.1, Fable 5, Opus 5.5, Opus 5 and Sonnet 5.5 carry safety classifiers that return an HTTP 200 response with stop_reason set to refusal and a stop_details.category such as cyber. The docs state plainly that benign cybersecurity work can also trigger the cyber category. A cyber refusal before any output is not billed, and an optional server-side fallback (beta) can retry a refused request on a model Anthropic recommends for that category, which may be less capable.

Anthropic's announcement says the Fable 5.1 classifiers block penetration testing, exploit generation, binary-based vulnerability scanning and malicious agentic coding, while allowing defensive identification of vulnerabilities, and that Claude Code users should see around 60 percent fewer cyber interventions per session than under Fable 5's safeguards. Its Cyber Verification Program gives verified users certain Opus and Sonnet class models with reduced cyber safeguards, with Mythos class models described as coming soon.

OpenAI's verification is the heaviest of the three as described. Its documentation says Daybreak access is controlled through identity verification, account security, monitoring, approved-use restrictions and legal attestations, that individuals apply through ChatGPT and organisations through a form, and that applying does not guarantee approval. Forkast reports that since 1 September 2026 every Daybreak Red account must also have a hardware security key; we could not confirm that on an OpenAI page. On the safeguard side, OpenAI's GPT-6 Astra card describes monitoring of complete trajectories including chains of thought, with flagged items scored for human review, and OpenAI's docs say a safeguard can block, reroute or limit a request.

Google's Fairwind page asks partners to restrict access to internal security, incident response or penetration testing teams, to use user-level authentication with phishing-resistant MFA, and to use the models only for authorised defensive and academic work. Priority goes to governments and national cyber authorities, critical infrastructure operators and core technology platforms. Google says Argon's general release will refuse harmful requests while preserving dual-use research under its Frontier Safety Framework, with monitoring of internal activations to spot misuse and of the model's chain of thought and actions for misalignment.

How to use it

Most readers meet this trend from the general-release side: a request about a CVE, a fuzzing harness or a log full of exploit strings comes back refused. You need nothing new to handle that. To get the relaxed tier you need to be doing authorised security work, usually inside an organisation, and to pass a review that none of the three labs promises to complete on a timeline.

  1. Find out which of your workloads are security-adjacent: code review of auth paths, dependency triage, malware or log analysis, pen-test tooling, CTF material. These are the calls most likely to hit a cyber classifier.
  2. On the Claude API, branch on stop_reason equal to refusal and log stop_details.category. Do not parse the explanation text; Anthropic says it is not stable.
  3. Decide your fallback deliberately. Anthropic's server-side fallback (beta header server-side-fallback-2026-07-01) retries on a recommended model and names the model that answered, which may be less capable. On OpenAI, check client notices and request logs, since its docs say a request can be rerouted rather than refused.
  4. If your team does defensive security professionally, apply to the programme that matches your stack: Anthropic's Cyber Verification Program, or Project Glasswing through your Anthropic, AWS or Google Cloud account team; OpenAI Daybreak through ChatGPT or OpenAI's enterprise form; Google's Fairwind Program through its application form.
  5. For OpenAI models on Amazon Bedrock, AWS's 11 August post says you first enrol in Trusted Access for Cyber, then request access through your AWS account team, and that the Daybreak models were in US East (Ohio) only at that date.
  6. Prepare the controls the programmes ask for before you apply: named users, phishing-resistant MFA or hardware keys, access limited to the security team, and an audit trail of what the model was asked to do.
  7. If you use Codex, OpenAI's DevDay update (29 September) puts Daybreak Blue models inside Codex Security Cloud without a separate Daybreak application, as reported. That is the shortest route to a relaxed tier we found.

Use cases

Sources

  1. Gemini 4 Argon, Google, 2026-09-30
  2. Fairwind Program, Google DeepMind, 2026-09
  3. Claude Fable 5.1 overview, Anthropic docs, 2026-09-01
  4. Claude Mythos 5.1 overview, Anthropic docs, 2026-09-01
  5. Claude Fable 5.1 and Claude Mythos 5.1, Anthropic, 2026-09-01
  6. Refusals and fallback, Anthropic docs, 2026-09
  7. Project Glasswing, Anthropic, 2026-04-07
  8. GPT-6 Astra, OpenAI Deployment Safety Hub, 2026-09-03
  9. Addendum to GPT-6 Astra System Card: GPT-6.1 Sol, Trusted Access for Cyber, OpenAI Deployment Safety Hub, 2026-09-29
  10. Models and Trusted Access, OpenAI (ChatGPT docs), 2026-10
  11. Daybreak Red and Daybreak Blue now available to eligible customers on Amazon Bedrock, AWS, 2026-08-11
  12. Google releases Gemini 4 Argon, called its most powerful model yet, TechCrunch, 2026-09-30
  13. Gemini 4 Argon: Google's new flagship reaches cyber defenders first, The Next Web, 2026-09-30
  14. GPT-6 Astra is the first model OpenAI classifies as Critical for cybersecurity, InfoQ, 2026-09-17
  15. OpenAI launches GPT-5.6-Cyber with reduced safeguards for exploit development, The Hacker News, 2026-08-11
  16. OpenAI to unveil GPT-6 Cyber model and a cybersecurity-focused product, Fortune, 2026-09-24
  17. OpenAI builds new security gateway to deploy GPT-6 Cyber, PYMNTS, 2026-09-25
  18. OpenAI's fourth cybersecurity model in twelve months is about gated access to dangerous capabilities, Forkast, 2026-09-25
  19. OpenAI expands Codex and its API at DevDay with security scans, a Decisions API, and Ultrafast, The Decoder, 2026-09-29
  20. OpenAI DevDay 2026, BenchLM, 2026-09-29