
Here's a strange kind of tech story: a company builds something remarkable, tells you almost everything about it, then tells you that you probably can't use it. That's GPT-5.6-Cyber in one sentence. OpenAI's cybersecurity model completes 95% of prompts asking it to build exploit chains, escalate privileges, or bypass authentication. Its general-purpose sibling, GPT-5.6-Sol, completes the same category of request just 1.5% of the time. Unless you work security at a large company, a security vendor, or one of a specific list of under-resourced defenders, you're not getting an account.
One quick note before anything else. Frontier AI labs, OpenAI included, ship new models, access tiers, and pricing changes on a timescale measured in weeks, not years. Treat everything below as accurate as of September 8, 2026, and check OpenAI's own Daybreak pages directly before assuming anything here still holds.
This isn't a "should you sign up" review, because for almost everyone reading it, signing up isn't on the table. It's the more interesting question underneath: what did OpenAI build, why is it locked up this tightly, who is the vetting bar really built for, and what does a model like this existing mean for the rest of us who'll only ever read about it.
GPT-5.6-Cyber at a Glance
Here's the whole shape of it in one table, before we get into what any of these numbers actually mean.
Spec | Detail |
|---|---|
Maker | OpenAI |
Announced | August 11, 2026 |
Built on | GPT-5.6-Sol |
Model ID | gpt-5.6-cyber |
Context window | Up to 400,000 tokens (long-context pricing applies past 272,000) |
Max output | 128,000 tokens |
Input price | $12.50 per million tokens, standard, OpenAI API |
Output price | $75 per million tokens, standard, OpenAI API |
Input types | Text and image |
Output types | Text only |
Availability | Daybreak Red only. Application, identity verification, and OpenAI approval required. No self-serve signup. |
Notice that last row. Everything above it reads like a normal model spec sheet. That last row is the entire story of this post.
What Is GPT-5.6-Cyber, and Why Does It Exist?
GPT-5.6-Cyber isn't built from scratch. It's GPT-5.6-Sol, OpenAI's mainstream flagship at the time, retrained for one purpose: advanced, authorized cybersecurity work, specifically locating zero-day vulnerabilities and building the exploit chains that prove they're real. OpenAI announced it on August 11, 2026, inside a post titled "Expanding Daybreak as the Cyber Defense Window Narrows." The argument: attackers are already using AI to move faster, and defenders have a shrinking window to close the gap before that becomes the norm.
It isn't OpenAI's first cyber-specific model either. A predecessor, GPT-5.5-Cyber, came before it, and the completion-rate numbers side by side show how fast this category is moving: 57.3% on GPT-5.5-Cyber's version of the evaluation, 95% on GPT-5.6-Cyber's, a 38-point jump in one generation on the exact test built to measure how often a model does the work rather than declines it.
One more sign of how new this is: OpenAI hasn't published a full system card for GPT-5.6-Cyber yet, the safety writeup it normally releases alongside a launch, saying one is coming "at a later date." For a model built to do things a model is normally trained to refuse, being asked to trust the vetting process before the safety paperwork exists is worth noticing.
What's Actually New: A Model Built to Say Yes
Most AI safety stories from the last few years are about models learning to say no more reliably. GPT-5.6-Cyber runs the opposite exercise, on purpose. OpenAI built it with what it calls a lower refusal rate for dual-use tasks, security work that helps defenders and helps attackers depending entirely on who's asking and why. That isn't a side effect of training, it's the entire point of the model, and the entire reason it's locked behind a vetting process instead of a signup form.
The clearest way to see it is the completion-rate gap on prompts covering exploit chain development, privilege escalation, and authentication bypass: GPT-5.6-Sol, the same base model with its normal guardrails intact, completes the request 1.5% of the time. GPT-5.6-Cyber completes the same category 95% of the time. A general-purpose model refuses this work almost every time it's asked. Its specialized sibling, built from the same weights, does it almost every time instead.
Worth saying plainly, since it affects how much weight to put on it: that 95% versus 1.5% versus 57.3% comparison, first reported in detail by SecurityWeek, is OpenAI's own internal evaluation, not a published third-party benchmark with an independent name attached, the way ARC-AGI or SWE-bench are, and no outside organization has published an independent replication of it. That doesn't make the number meaningless, a 95-to-1.5 gap on a company's own test is still a real, deliberate design choice, reported honestly. It just means the right way to read it is "OpenAI says," not "independently confirmed."
What the Benchmarks Actually Say
Zoom in past the headline number and the picture gets more textured, and more honest, than a single completion rate suggests.
Start with a genuine result: in OpenAI's own evaluation, GPT-5.6-Cyber found 400 privilege-escalation vulnerabilities in an operating system kernel during testing, the kind of discovery that would take a human researcher a long time to match. But even OpenAI's own framing carries an honest catch: finding 400 vulnerabilities is discovery, not triage. Someone still has to verify, prioritize, and patch all 400, and a model extremely good at finding problems isn't automatically as good at telling you which ones matter most. Discovery power without triage capacity isn't a finished tool, it's a large backlog with a head start.
That tension shows up again in a comparison OpenAI ran between Cyber and Sol accessed through Daybreak Blue, on the same tasks. At a standard setting capped at 300 turns, Sol through Daybreak Blue actually solves tasks more efficiently, using fewer tokens, with better-documented output. Cyber sometimes produces shorter, less detailed vulnerability reports on the same task, a real cost when a human has to act on the report afterward. Extend the test to 600 turns and the gap narrows but doesn't close. The honest reading: Cyber trades some report quality and token efficiency for its willingness to engage with harder, higher-risk requests. That's a real trade, not a strictly better version of Sol.
How much of this has anyone other than OpenAI checked? No fully independent evaluation of GPT-5.6-Cyber specifically has been published. The closest adjacent data point is Irregular, an outside AI security lab that evaluated Sol, not Cyber, on real-world offensive tasks and found modest gains over the prior generation. Even OpenAI's outside evaluators for the broader GPT-5.6 family work under real constraints: METR's predeployment assessment of Sol was conducted under NDA, with OpenAI's own legal and communications teams reviewing the writeup before it went public. That's becoming standard across the industry, not unique to OpenAI, but "OpenAI's evaluator confirmed it" and "an independent party confirmed it with no vendor involvement" aren't quite the same sentence. That distinction deserves more scrutiny than it's had, and the honest answer so far is that it hasn't gotten it yet.
What GPT-5.6-Cyber Actually Costs
Here's a wrinkle worth flagging plainly, because the assumption most coverage runs with is that GPT-5.6-Cyber simply has no public price tag. That's true in one specific place: OpenAI's consumer-facing pricing page lists Sol, Terra, and Luna with real numbers and leaves Cyber's row blank. But the model's technical documentation, and its Amazon Bedrock listing, both carry actual figures: $12.50 per million input tokens, $1.25 per million cached input, and $75 per million output on OpenAI's own infrastructure; a marked-up $13.75, $1.375, and $82.50 through Bedrock. Cross the 272,000-token context threshold and, matching the rest of the GPT-5.6 family, the entire request re-rates at the higher rate, not just the tokens past the line.
So "no public pricing" is narrower than it sounds. The number is filed in documentation for people already partway through vetting, not on a page anyone can browse. For almost everyone reading this the distinction is academic anyway, since you can't spend that money without clearing Daybreak Red's approval first.
For context, GPT-5.6-Sol itself currently runs $4 per million input tokens and $20 per million output tokens, a promotional rate through November 21, 2026, putting Cyber at roughly three times Sol's price on top of an approval most readers won't get. And plenty of real-world GPT-5.6-Cyber usage doesn't flow through a direct API key at all: OpenAI's Daybreak Cyber Partner Program routes access through vendors like Palo Alto Networks, CrowdStrike, and Sophos, who build the model into their own products while access stays with the partner. Most organizations that benefit from GPT-5.6-Cyber will do it by buying a security vendor's product, not by opening an OpenAI billing dashboard.
Where It's Strong, Where It Falls Short
Where it's strong:
A 95% completion rate on exploit chain, privilege escalation, and authentication bypass prompts, against 1.5% for unmodified GPT-5.6-Sol and 57.3% for predecessor GPT-5.5-Cyber, per OpenAI's own internal evaluation.
Found 400 privilege-escalation vulnerabilities in an OS kernel during OpenAI's own testing, a genuinely large discovery number.
Built on the same base as GPT-5.6-Sol, so it inherits a capable general model rather than being a narrow bolt-on tool.
A real distribution path: partners including Accenture, IBM, Palo Alto Networks, CrowdStrike, and Cloudflare are already building it into shipping products.
Where it falls short:
No fully independent, non-vendor evaluation of GPT-5.6-Cyber itself has been published as of this writing.
Produces shorter, less detailed vulnerability reports than Sol on the same tasks at standard settings, and uses more tokens getting there.
No system card published yet; OpenAI says one is coming later.
Text and image input only, no audio or video, limiting its use in multimodal "computer use" security workflows.
Essentially unavailable to individual developers, freelancers, or small teams without an enterprise relationship: no public pricing page, no self-serve signup.
Daybreak Blue vs Daybreak Red: Who Actually Gets Through
Daybreak is the program that decides who gets any of this, split into two tiers solving different problems.
Daybreak Blue is the wider door: approved users get GPT-5.6-Sol and other general-purpose models with guardrails recalibrated for defensive work, vulnerability discovery, malware analysis, incident response, patch validation, so the model doesn't reflexively refuse legitimate security questions. OpenAI positions it as the starting point most security teams should use. Daybreak Red is the narrow door that includes GPT-5.6-Cyber specifically, and OpenAI is blunt about who qualifies: an organization has to document real authorization for the work, constrain the model's system access, and actually review what it does. That's the practical bar, not a marketing line: provable permission to test the systems in question, an access scope the model can't wander outside of, and a human checking the output rather than trusting it as a black box. The tiers aren't a ladder either, approval for Blue doesn't put you any closer to Red; each needs its own application and verification.
For an individual, the baseline is simple to state and hard to clear: at least 18 years old, then evaluated on identity and trust verification, risk considerations, the intended use case, and, notably, your ability to strengthen the broader cybersecurity ecosystem, not just your own environment. It isn't enough to have a legitimate use case for yourself; OpenAI is explicitly weighing whether approving you makes defenders collectively better off.
The partner list is the more honest signal, though. Confirmed Daybreak Cyber Partners include Accenture, Capgemini, EY, IBM, KPMG, and PwC on the consulting side, and Palo Alto Networks, Sophos, CrowdStrike, Fortinet, Akamai, and Cloudflare on the security vendor side. Read that list for what it is: every household name in enterprise security and Big Four consulting, and basically nobody else. That's not a knock on the program, an exploit-capable model genuinely needs a high bar. But it does mean the honest starting assumption for a solo developer or small team is that you're not the intended user, at least not yet, and not through this door.
The $1 Billion Program for Defenders Without Enterprise Budgets
Here's the part of this story that actually matters if you're not Accenture or CrowdStrike.
On September 4, 2026, less than a month after GPT-5.6-Cyber's launch, OpenAI committed $1 billion to subsidize Daybreak access, training, and technical support for a specific list of organizations that tend to defend critical infrastructure on nonexistent security budgets: water and wastewater systems, electric grid operators, state and local governments, community and regional banks, nonprofits, and open-source projects. That list reads like a rundown of exactly who gets targeted constantly and can least afford a security team, presumably the point.
The mechanics matter too. The billion isn't cash, it's credit against OpenAI's own products, targeted to be consumed within six months, starting in the United States with plans to extend to partner countries within weeks. A pilot is already running with the Multi-State Information Sharing and Analysis Center, a federally supported group that shares threat intelligence with public-sector organizations, and OpenAI says it convened utility representatives from 40 states and the District of Columbia, covering essential services for more than half the US population, ahead of the announcement.
If you run security, even informally, for a nonprofit, an open-source project, a community bank, or a local government office, this is the realistic path worth looking into, more realistic than applying to Daybreak Red cold. That said, be clear-eyed about how thin the details still are: OpenAI hasn't published specific cost figures or a detailed eligibility checklist beyond the MS-ISAC pilot, which matters if you're a small county IT office trying to work out whether you qualify. Applying and finding out is still the only real way to know.
Why This Still Matters If You'll Never Touch It
Step back and there are two separate stories happening in OpenAI's lineup right now, worth telling apart. GPT-6 Astra, released September 3, 2026, is a general-purpose flagship that turned out capable enough at security work, it found two previously unknown vulnerabilities in V8, Chrome's JavaScript engine, that OpenAI classified it at the "Critical" cybersecurity capability tier under its own Preparedness Framework, with standard access refusing advanced offensive tasks by default, pending wider Daybreak access. GPT-5.6-Cyber is the opposite kind of model: built deliberately to do what a general model refuses, gated from day one rather than restricted after the fact. One is a general model that got too capable to leave unguarded. The other is a specialist built with guardrails intentionally lowered for people already vetted. Different problem, same underlying reason: this category of capability has outgrown "just let anyone use it."
There's a related, smaller effort worth knowing about too. Aardvark, OpenAI's agentic security researcher powered by GPT-5, has been running in private beta, scanning codebases to find, validate, and help patch vulnerabilities continuously rather than on a schedule. Participants get early access and work directly with OpenAI to refine detection accuracy, validation, and reporting, a lower-stakes preview of the same idea driving Daybreak: AI is now good enough at parts of security work that access itself is the thing worth managing carefully.
Here's the tension worth sitting with honestly, since it doesn't resolve cleanly either way. Lowering a model's refusal rate for offensive security work genuinely helps defenders, chronically outnumbered against attackers who don't need anyone's permission to experiment. That's a real case for what OpenAI built. It's also true that doing so concentrates real capability in whoever passes the vetting, and vetting is run by companies, not neutral referees, however carefully designed. Reasonable security professionals land in different places on whether gated access is the right tradeoff or just a more comfortable-looking version of the same risk. The vetting bar itself, document authorization, constrain system access, review the model's actions, plus a partner list that's almost entirely Fortune 500 consultancies and security vendors, is the actual evidence of where OpenAI is drawing that line right now. Whether that's the right place is a genuinely open question worth reaching your own conclusion on, rather than taking OpenAI's framing at face value.
Who Should Use It, Who Should Skip It
Use, or at least apply for, GPT-5.6-Cyber if you run security for an organization that can document real, specific authorization for offensive testing work, keep the model's access constrained to that scope, and put a human in the loop reviewing everything it produces, particularly at a security vendor, a large consultancy, or an organization the $1 billion Frontline Defenders program is explicitly built for.
Skip it, at least as a direct signup target, if you're a solo developer, a freelancer, or a small team without a documented, authorized security engagement to point to. That's not a judgment on your work, just an accurate read of who Daybreak Red is built to approve right now. For sensitive code review or security-adjacent work you'd rather not run through any vetted-access program, a self-hosted model through a tool like Ollama keeps everything on your own hardware, real limits on raw capability but zero gatekeeping. And for the far more common problem most readers actually have, catching bugs before they ship, a mainstream tool like CodeRabbit solves a version of the same problem GPT-5.6-Cyber solves for defenders, aimed at ordinary bugs instead of zero-days, and available to anyone with a GitHub repo.
Verdict: The Model You Probably Can't Have
Here's the honest verdict, a different shape than most posts on this site: GPT-5.6-Cyber is a real, meaningful capability jump, verified by OpenAI's own numbers if not yet by anyone outside OpenAI, and it exists specifically because defenders needed something a general-purpose model was built to refuse them. It's also, for the overwhelming majority of people who'll read about it, not a tool you can sign up for right now. No pricing page to browse, no free trial, no waitlist that resolves into an account through persistence alone.
If that describes you, the useful takeaway isn't "try it anyway." It's this: if you work security for critical infrastructure, a nonprofit, or an open-source project, the $1 billion Frontline Defenders program is genuinely worth applying to, more so than Daybreak Red cold. If you're a solo developer or small team who wants better AI help with code and security without anyone's vetting process, Ollama and CodeRabbit already solve real, adjacent problems, with zero gatekeeping. And if you're just deciding which everyday AI assistant to use for the other 95% of your work, that's a simpler, unrelated decision we've covered in our ChatGPT vs Claude comparison.
What GPT-5.6-Cyber mostly tells you isn't really about the model itself. It's that the industry has quietly reached a point where "how capable can we make this" and "who should be allowed to use it" are now two separate, equally important product decisions, and the second one just became a whole program with its own tiers, vetting criteria, and billion-dollar subsidy. That's worth understanding even from outside the velvet rope.
Curious to see how it performs?
Try
ChatGPT
Now
GOT ANY QUESTIONS LEFT?
What is GPT-5.6-Cyber?
How do I get access to GPT-5.6-Cyber?
What is the difference between Daybreak Blue and Daybreak Red?
Is GPT-5.6-Cyber safe to release?


