Ask a frontier AI model to help hunt for a software vulnerability, and there’s a good chance it refuses. Security teams have spent months fighting that reflex, watching legitimate penetration tests get flagged as attacks by the same guardrails meant to stop hackers.

This has become a running frustration inside corporate security teams that are otherwise desperate for more firepower. They need this firepower because threat actors are improving their skills everyday, looking for ways to exploit weaknesses.

OpenAI’s answer, unveiled on Monday, Aug. 10, says almost as much about the risk sitting inside its own models as it does about defense.

OpenAI is expanding its Daybreak cybersecurity program into two access tiers, the company said. Daybreak Blue strips the cyber-related safety filters off GPT-5.6 Sol, the company’s flagship model, for vetted defenders doing everyday security work.

Daybreak Red goes further, granting access to a new model called GPT-5.6-Cyber, built specifically for exploit validation and advanced vulnerability research.

The company frames the split as narrowing the gap between attackers and defenders, warning that hackers will increasingly use AI to launch attacks at machine speed.

GPT-5.6-Cyber is built on GPT-5.6 Sol but trained to reduce refusals on cybersecurity tasks that would otherwise trip its safety filters, OpenAI said.

OpenAI’s Daybreak Blue and Red split access, not intent

The distinction between Blue and Red is not really about who can be trusted. Both tiers require identity verification and legal attestations, and starting Sept. 1, individual accounts must adopt hardware security keys. The real distinction is how close a user can get to raw offensive capability.

Daybreak Blue is the tier OpenAI recommends for most organizations, built for day-to-day defensive work. Daybreak Red is reserved for experienced defenders tackling harder problems, the kind of work that blurs into the exact skills an attacker would need.

Both tiers pull from the same underlying model family, which means the real product OpenAI is selling is trust, not technology.

OpenAI splits its Daybreak cybersecurity program into two tiers, days after regulators flagged its own flagship model for unauthorized activity.

ALEX WROBLEWSKI / Getty Images

The real gate is a capability threshold

The more revealing decision sits next to the launch. OpenAI said last week it is pausing some internal work on a more advanced model called Astra after it showed significant advances in agentic coding and cybersecurity. GPT-5.6-Cyber shipped anyway, days later.

That pairing suggests OpenAI is gating releases by what a model can do, not by who is asking to use it. A model capable enough gets held back, regardless of the access controls wrapped around it.

One that falls just short of that line ships instead, guarded by identity checks rather than a pause.

OpenAI’s own models keep going rogue

The urgency behind Daybreak has a direct cause. In July, an OpenAI model broke out of a testing environment and accessed the AI platform Hugging Face without authorization, an incident OpenAI itself described as a turning point for the industry.

Weeks later, the U.K.’s AI Security Institute found that both GPT-5.6 Sol and Anthropic’s Mythos 5 model engaged in sustained, potentially harmful activity against real organizations during separate evaluations, according to Bloomberg.

More OpenAI:

That finding complicates OpenAI’s own pitch. GPT-5.6 Sol is the same model Daybreak Blue hands to vetted defenders with its guardrails stripped away, and regulators had already flagged it for acting outside its intended bounds.

Anthropic disclosed a similar pattern days earlier, saying its Claude model breached three organizations after slipping past a sandbox meant to keep it offline, according to Bloomberg.

The incidents have already reached Congress. More than 1,000 employees across OpenAI, Anthropic, and other labs signed an open letter last month urging the government to help pace the speed of AI development, according to CBS News.

Lawmakers introduced legislation that would require AI companies to maintain the ability to shut down or throttle their models, a response that directly referenced the Hugging Face breach.

A pattern now repeating across every lab

OpenAI is not the first lab to build a gated tier around its most capable cyber model. Anthropic launched Project Glasswing months earlier, giving 12 partner organizations early access to a cybersecurity-focused preview of its Mythos model.

Anthropic framed the coalition as a race to secure critical software before comparably powerful cyber models from OpenAI and Google reached wider release, Fortune reported. Daybreak is OpenAI’s answer to a structure its rival built first.

Both companies are converging on the same uncomfortable conclusion. The skills that make a model good at finding vulnerabilities are the same skills that make it dangerous, and no amount of vetting fully separates the two.

Security leaders are already blending frontier models with open-source tools rather than betting on any single lab’s access controls. That hedge says more about where confidence in AI safety programs actually stands than any tier name or benchmark score.

As more labs release cyber models with fewer guardrails, the real test will not be which company builds the smartest defender. It will be which one is first to prove its own model cannot be turned against the people using it.

Related: Tech expert predicts an OpenAI collapse