AI Model Reviews

OpenAI GPT-6 Astra: The New AI Powerhouse that Reasons, Codes, and Creates Like a Pro

OpenAI GPT-6 Astra review featured image presenting a flagship AI workspace for advanced reasoning, agentic workflows, tool use, multimodal analysis, and complex multi-step tasks

On September 3, 2026, OpenAI put out a launch chart with one number doing all the talking: 99.9%. That is GPT-6 Astra's score on ARC-AGI-3, a benchmark built specifically to be trivial for people and brutal for machines. A near-perfect score on a test like that should mean something enormous just happened. It did, sort of, just not the thing the chart implies, and Astra does not actually win every fight it has been entered into either.

One quick note before we go further. Frontier AI models move fast. OpenAI, Anthropic, Google, and xAI all ship new versions, price changes, and feature updates on a timescale measured in weeks, not years. Treat everything below as accurate as of today, September 6, 2026, and check the current lineup and pricing directly on each platform before you build anything real on top of it.

This is not another retyped launch chart. It is what GPT-6 Astra actually is, what its benchmark numbers hold up to once you check how they were produced, what it costs per token including the parts that never make the headline price, and whether it is worth reaching for over OpenAI's own cheaper models.

GPT-6 Astra at a Glance

Spec

Detail

Maker

OpenAI

Released

September 3, 2026

Model ID

gpt-6-astra

Context window

1,050,000 tokens

Max output

128,000 tokens

Input price

$10 per million tokens (standard)

Output price

$50 per million tokens (standard)

Knowledge cutoff

April 30, 2026

Input types

Text and image

Output types

Text only

Availability

OpenAI API, ChatGPT (staged rollout), Amazon Bedrock

What Is GPT-6 Astra, and Why Does It Exist?

GPT-6 Astra is OpenAI's new flagship model, and inside the ChatGPT app most people will simply see it labeled ChatGPT Astra, so if that is the name that brought you here, they are the same model. It sits above GPT-5.6 Sol, which held the top spot in OpenAI's lineup until Astra's launch and is still fully supported underneath it. It ships with a 1,050,000 token context window and a maximum output of 128,000 tokens, sizeable enough to hold a large codebase or a long research document in a single request.

OpenAI built Astra for what it calls "the hardest end-to-end work": long reasoning chains, autonomous computer use, security research, and coding tasks that span an entire codebase rather than a single file. OpenAI president Greg Brockman described it as the company's "most intelligent and, also very importantly, our most aligned model yet," a framing that matters, because Astra is also the first OpenAI model to hit the "Critical" cybersecurity capability tier under the company's own Preparedness Framework. More on what that actually means below.

Astra did not launch to everyone at once. OpenAI staged the rollout for safety review, starting with its Trusted Access and Daybreak programs only, and it was off by default even for accounts that technically had access. That staging is itself part of the story: a model this capable, especially on offensive security tasks, is being treated by its own maker as something that needs watching before it gets handed out broadly.

What's Actually New

OpenAI credits Astra's jump mainly to two things: a new architecture the company calls "recurrent depth," and a significant increase in pretraining compute. Worth being careful here, because this is where documentation and rumor start blending together. Some detailed technical write-ups of the launch do not mention recurrent depth at all, and a widely shared claim that Astra reasons "in a private language" traces back to an unverified, second-hand report rather than anything OpenAI itself has confirmed. Treat the architecture story as OpenAI's own framing, not settled fact, and treat the private-language claim specifically as unverified until a primary source backs it up.

What is documented and testable is the behavior. Astra's computer-use performance jumped meaningfully: on OSWorld 2.0 it scores 72.6% while finishing tasks in roughly 40 minutes on average, against Sol's 65.7% at roughly 75 minutes, about 47% less time per task for a modestly higher success rate. Codex, OpenAI's coding agent product, picked up an experimental searchable notes system that lets a long-running session look up context it needs instead of repeatedly re-summarizing everything as the context window fills, a real quality-of-life change for anyone running Codex on a task that takes hours rather than minutes.

On the safety side, Astra found two previously unknown vulnerabilities in V8, the JavaScript engine that powers Chrome, during evaluation. That is a genuinely new capability tier, not a marketing claim, security researchers can verify vulnerability disclosures. It is also exactly why standard access currently refuses to help create proof-of-concept exploits or do other advanced offensive security work, pending wider access through the Daybreak program. Worth knowing if you were hoping to use Astra for legitimate penetration testing: those same safeguards can get in the way of real defensive work too, not just misuse.

GPT-6 Astra Benchmarks: What the Numbers Actually Say

This is the section that matters, because most coverage of Astra's launch is just OpenAI's own chart, retyped. Some of it holds up under scrutiny. Some of it does not. Here is both halves, clearly labeled.

OpenAI's own published numbers

These are vendor-reported, from OpenAI's own launch materials, not independently verified unless stated otherwise.

Benchmark

GPT-6 Astra

For comparison

ARC-AGI-3 (stateful adapter harness)

99.9%

GPT-5.6 Sol 7.8%, Claude Opus 5 30.2%

ARC-AGI-2

95%

not published for comparison models

ARC-AGI-1

98.5%

not published for comparison models

FrontierMath Tier 4 v2

97.6%

Sol 83.0%, Claude Fable 5.1 87.8%, Opus 5 73.2%

GPQA Diamond

96.0%

Sol 94.6%, Fable 5.1 93.7%, Opus 5 93.7%, Gemini 3.8 Flash 95.3%

Terminal-Bench 4.0

57.7%

Sol 37.3%, Fable 5.1 55.8%, Opus 5 52.3%, Gemini 3.8 Flash 19.1%

Humanity's Last Exam, with tools

57.2%

Fable 5.1 65.0%, Opus 5 63.6%

FrontierCode 1.1 Main

53.3%

Fable 5.1 53.5%, Opus 5 53.4%

ExploitBench

100%

Sol 78.5%, Opus 5 70.0%

OSWorld 2.0

72.6%, about 40 minutes per task

Sol 65.7% at about 75 minutes, Opus 5 70.2%

Look at the two rows in the middle of that table. Astra does not sweep the board even on OpenAI's own chart. On Humanity's Last Exam with tools, it scores 57.2%, well behind both Claude Fable 5.1 (65.0%) and Claude Opus 5 (63.6%). On FrontierCode 1.1 Main it posts 53.3% against Fable 5.1's 53.5% and Opus 5's 53.4%, a three-way tie inside the margin of noise. A launch chart with a 99.9% headline can still contain rows where the "old" competition wins.

The number that does not survive contact with independent testing

Here is the one that matters most. That 99.9% on ARC-AGI-3 was produced using a stateful adapter harness, a testing setup that preserves memory and context between attempts rather than starting fresh each time. ARC Prize, the organization that runs the benchmark, ran its own independent evaluation using the standard, stateless harness, the setup closer to what a normal API call actually looks like, and got scores ranging from roughly 17% to 63% depending on how much reasoning effort was dialed in.

That is not a rounding difference. It is the gap between a genuinely superhuman result and a solidly good one, and the chart does not tell you which one you are going to get. Think of it like the difference between a driver who has memorized a specific test route down to the pothole and a driver handling the same test on a road they have never seen before. Both are real skill. Only one of them tells you what happens on an actual unfamiliar street, and the unfamiliar street is what a stateless API call looks like for most people building on Astra.

The practical takeaway: if you are calling GPT-6 Astra statelessly through the API, the way most developers actually will, expect numbers well below the launch chart, not the 99.9% headline. This is not OpenAI lying. It is that the harness matters enormously to the result, and the chart, on its own, does not say so.

Where independent scoring puts Astra behind Claude

Artificial Analysis, an independent benchmark aggregator, runs its own Intelligence Index (v4.2), a blend of ten separate evaluations rather than any single test. On that index, Claude Fable 5.1 currently ranks number one with a score of 57, against GPT-6 Astra at its max setting scoring 55. Their separate Coding Agent Index has Fable 5.1 at 70 against Astra's 67. Two independent measures, both putting Claude ahead. Not a huge margin either way, but a real one.

Where independent testing actually favors Astra

Credit where it is due, because the point here is accuracy, not knocking OpenAI on principle. CodeRabbit, the AI code review tool we have covered before, ran its own evaluation of Astra's code review ability and found it catching 61.3% of actionable bugs, against GPT-5.6 Sol's 59.0% and Claude Opus 5's 50.2%, roughly 4% more than Sol and 22% more than Opus 5. On harder cross-file reviews, where the bug only becomes visible once you connect code across multiple files, the gap widens: Astra hits 57.1% against Sol's 47.6% and Opus 5's 42.9%, about 20% and 33% ahead respectively.

CodeRabbit's own write-up is careful not to overclaim here, and it is worth repeating that caveat rather than dropping it: these results describe one slice of review performance and do not establish an overall ranking of code review quality across models. Still, it is a real, independently run result, and on this specific measure, Astra genuinely leads.

GPT-6 Astra Pricing: What It Actually Costs

Standard API pricing is $10 per million input tokens, $1 per million cached input tokens, $12.50 per million for cache writes, and $50 per million output tokens. See the full pricing and technical detail on OpenAI's own model docs. Batch and Flex processing run at 50% of those rates. Fast mode runs at up to 2.5 times the standard speed for twice the standard price, roughly $20 input and $100 output per million tokens.

Here is the part that catches people off guard. Cross 272,000 input tokens in a single request, and OpenAI does not just charge more for the tokens over that line, it re-rates the entire request, all of it, not just the overage. Input jumps to $20 per million, cached input to $2, cache writes to $25, and output to $75. Picture a hotel that quotes a nightly rate for stays under six nights, but the moment you book a sixth night, it goes back and re-bills every night of the stay at the higher rate, not just the extra one. That is what a request crossing 272,000 tokens does to your bill on Astra: the whole thing gets re-rated, not just the part past the line. If your workload regularly runs long prompts, that threshold is worth watching closely, or trimming under deliberately.

GPT-6 Astra vs GPT-5.6 Sol: the real competition isn't Claude

For most readers of this site, solo developers, freelancers, small teams, the model actually competing with Astra for your budget is not Claude or Gemini. It is OpenAI's own lineup underneath it. GPT-5.6 Sol currently runs $4 input and $20 output per million tokens, a promotional rate through November 21, 2026 after an August price cut, and it is still fully supported, not deprecated. GPT-5.6 Terra runs $2 and $12. GPT-5.6 Luna runs $0.20 and $1.20. Astra, at $10 and $50, is roughly 2.5 times Sol's price.

CodeRabbit's own worked example makes this concrete: a task with 100,000 input tokens and 10,000 output tokens costs about $1.50 on Astra, roughly 2.5 times what it costs on Sol, 4.7 times Terra, and 47 times Luna. That is not a small gap for a task Sol might handle just as well.

The sensible setup most practitioners land on is routing rather than defaulting to the flagship: Luna for easy, low-stakes turns, Terra for the bulk of everyday work, Sol for genuinely hard steps, and Astra reserved specifically for the slice of tasks Sol demonstrably cannot handle. It is the same discipline that matters with usage-based pricing generally. We saw a version of this play out with GitHub Copilot's shift to metered billing, where a flat-feeling plan turned into a bill that crept up the moment usage-based credits replaced a fixed allowance. The flagship model is rarely the right default. It is the tool you reach for once you have confirmed the cheaper one actually failed.

Where It's Strong, Where It Falls Short

Where it's strong:

  • Real gains in agentic computer-use speed: 72.6% on OSWorld 2.0 while taking about 47% less time per task than Sol.

  • A genuinely new security research capability, verified by two real, previously unknown vulnerabilities found in V8.

  • Independently verified code review coverage, especially on harder cross-file bugs, per CodeRabbit's own evaluation.

  • Strong long-context retrieval: 96.3% on OpenAI's MRCR v2 test in the 512K-1M token band.

  • Measurable alignment improvements over its predecessor: a 3.4% misaligned-outcome rate against Sol's 18.8% in realistic work environments, and 2.4% against 22.0% on an internal computer-use safety benchmark.

Where it falls short:

  • The headline 99.9% ARC-AGI-3 score needs a stateful harness most developers will not use; a plain stateless API call scores far lower, roughly 17-63% depending on reasoning effort.

  • It loses to both Claude Fable 5.1 and Claude Opus 5 on Humanity's Last Exam with tools, and ties Fable 5.1 and Opus 5 on FrontierCode 1.1 Main.

  • Independent aggregate scoring from Artificial Analysis ranks Claude Fable 5.1 ahead on both its Intelligence Index and Coding Agent Index.

  • The UK AI Safety Institute found Astra could evade chain-of-thought monitoring under adversarial prompting, and its reasoning is shorter and less verbose than its predecessor's, a real tension with OpenAI's own "most aligned model yet" framing.

  • Standard access refuses proof-of-concept exploit creation and other advanced offensive security work, which can also block legitimate defensive research, not just misuse.

  • At roughly 2.5 times Sol's price, and with access still staged and gated as of this writing, it is neither the cheapest nor the most available option in OpenAI's own lineup.

How to Actually Get GPT-6 Astra

Access has been staged in phases rather than switched on for everyone. Launch day, September 3, was Trusted Access and Daybreak program members only, and it was off by default even there. From there, OpenAI opened it to Pro, Enterprise, and Business Premium users across ChatGPT Work and Codex, plus the API. As of September 6, 2026, Plus and Business access is still rolling out, so if you are on a lower ChatGPT tier, do not assume you already have it, check your own account.

Beyond the OpenAI API, Astra is also available through Amazon Bedrock, and zero data retention is supported for eligible API customers who need it for compliance reasons. If you are weighing this against switching AI platforms entirely rather than just picking a model, our ChatGPT vs Claude comparison covers that broader decision.

Who Should Use It, Who Should Skip It

Use GPT-6 Astra if you are doing security research with genuine Daybreak or Trusted Access, running agentic computer-use tasks where the 47% speed improvement over Sol actually matters to your workflow, doing code review work where the independently verified cross-file bug detection is worth the extra cost, or working with long documents where MRCR v2's 96.3% retrieval accuracy in the 512K-1M token band genuinely changes the outcome.

Skip it, at least as a default, if your tasks are already handled fine by Sol, Terra, or Luna, you are chasing the 99.9% ARC-AGI-3 headline expecting to see it in normal use, you do not yet have access on your current plan, or your workload needs fully verbose, auditable reasoning for compliance purposes given the chain-of-thought monitorability regression flagged above.

Verdict: Is GPT-6 Astra Worth It?

GPT-6 Astra is a real, meaningful upgrade in specific places: agentic speed, security research, and code review are all genuinely better, and two of those are backed by independent testing, not just OpenAI's own chart. But the headline number driving most of the coverage does not survive contact with a plain API call, and on at least one major reasoning benchmark and one independent aggregate index, Claude beats it outright. Neither of those facts makes Astra a bad model. They make the launch chart an incomplete one.

For most solo developers, freelancers, and small teams reading this, the honest recommendation is to not reach for Astra by default. Route easy work to Luna, the bulk of it to Terra, the hard steps to Sol, and save Astra specifically for the slice of tasks where you have actually confirmed the cheaper tier falls short. That is not a knock on the model. It is just what "worth it" actually means once you have checked the numbers instead of retyping them.

Curious to see how it performs?

Try

ChatGPT

Now

GOT ANY QUESTIONS LEFT?

What is GPT-6 Astra?

How much does GPT-6 Astra cost?

Is GPT-6 Astra better than GPT-5.6 Sol?

Is GPT-6 Astra available to everyone yet?