
OpenAI's GPT-5.6 Luna costs 20 times less than the family's flagship, Sol, on input tokens, and on Agents' Last Exam, a benchmark built around long, professional-grade reasoning tasks, it scores 50.3 against Terra's 50.4. That is not a rounding error in Luna's favor. It is a dead heat with the tier one rung above it, priced at a fraction of the cost. Then you hand Luna a long document and ask it to find one specific fact buried inside, and it manages 41.3%. Terra manages 89.6% on the exact same test. That second number, not the price tag, is the thing you actually need to understand before you route production traffic to Luna.
Quick housekeeping before we go further: OpenAI ships new models, price cuts, and feature changes at a pace that makes any specific number here a moving target. Everything below is accurate as of September 8, 2026. Check developers.openai.com/api/docs/models/gpt-5.6-luna before you commit a production budget to it.
This post is about that split personality. Luna is not simply "the cheap, worse model" in OpenAI's lineup. It is a model that is nearly as smart as the tier above it, right up until your context gets long, and then it falls off a cliff. We will walk through the full three-tier pricing table (tabulated properly, in one place, instead of buried across separate docs pages), show what the benchmarks actually say without repeating OpenAI's own framing as if it were independent, and tell you plainly where Luna is the right call and where it will quietly cost you more than it saves.
GPT-5.6 Luna at a Glance
Spec | Detail |
|---|---|
Maker | OpenAI |
Family launch | July 9, 2026 (GPT-5.6: Sol, Terra, Luna) |
Price cut | Input and output prices cut about 80%, July 30, 2026 |
Model ID | gpt-5.6-luna |
Context window | 1,050,000 tokens |
Max input | 922,000 tokens |
Max output | 128,000 tokens |
Knowledge cutoff | February 16, 2026 |
Input types | Text and image |
Output types | Text only |
Endpoints | Chat Completions, Responses, Batch |
Input price | $0.20 per million tokens |
Output price | $1.20 per million tokens |
Availability | OpenAI API, ChatGPT Work, Codex |
That's the whole shape of the model in one table. The rest of this post is about the two numbers a spec table cannot show you: how close Luna actually gets to Terra, and exactly where it stops.
What GPT-5.6 Luna Is, and Why OpenAI Built It
GPT-5.6 launched on July 9, 2026, not as a single model but as a family of three: Sol, Terra, and Luna. Instead of shipping one flagship and calling it done, OpenAI split the release into what it calls durable capability tiers, each able to advance on its own cadence rather than forcing every task through the same heavyweight model.
Luna fills the role earlier "mini" and "nano" tier models played in OpenAI's lineup: the fast, cheap, high-volume option, built for the kind of prompt sent by the thousand rather than the one careful prompt run once. Think of it like this. Sol is the specialist you call in for the case that actually needs one. Terra is the capable generalist you default to. Luna is the model built for the request so routine that paying specialist rates for it would be silly, chat replies, tagging a support ticket, classifying a document, a lightweight agent step that just needs to move fast and cost little.
All three tiers share the same 1,050,000-token context window, the same 128,000-token max output, and the same February 16, 2026 knowledge cutoff. For the full breakdown of what else the three tiers share and exactly how they differ end to end, our GPT-5.6 Terra tier comparison covers that ground properly, we're only going deep on Luna here.
What's Actually New Since Launch
The single biggest change to Luna since it shipped isn't a capability upgrade, it's a repricing event, and it's worth being precise about that distinction. On July 30, 2026, OpenAI cut Luna's price by about 80%, from $1.00 input / $6.00 output per million tokens down to $0.20 / $1.20. Terra got a smaller cut the same day, about 20%, from $2.50 / $15.00 down to $2.00 / $12.00. Nothing about the underlying model changed that day, the weights are the same Luna that launched three weeks earlier, just priced very differently.
Two smaller but genuinely useful changes came with it. First, Luna and Terra usage inside ChatGPT Work and Codex subscriptions now consumes fewer credits than before, even though the subscription prices themselves didn't move. Second, OpenAI upgraded Auto-review in the ChatGPT app and in the Codex CLI from GPT-5.4 to GPT-5.6 Luna, a concrete, always-on production use, not just a bargain-bin fallback option nobody actually routes to.
One more thing worth flagging early, because it sets up the whole rest of this post: Luna shipped with the full 1,050,000-token context window, same as Terra and Sol, at the cheapest price in the family. A huge context window at a rock-bottom price sounds like an unambiguous win. Whether Luna can actually use that window is a separate question, and it's the one the benchmarks answer next.
What the Benchmarks Actually Say
Where Luna genuinely keeps up
On Agents' Last Exam, a benchmark built around long, professional-grade reasoning tasks across 55 fields, Luna scores 50.3, Terra scores 50.4, and Sol scores 53.6. That's on OpenAI's own published benchmark reporting, worth stating plainly since it's the vendor's own chart, not an outside lab's. Still, a 0.1-point gap between Luna and Terra is close enough to call a tie, and Terra costs ten times more per input token to run.
Independent testing tells a similar story. On the Artificial Analysis Coding Agent Index, a benchmark run by a third party rather than OpenAI itself, Luna scores 74.6, within striking distance of Terra's 77.4 and Sol's 80. Artificial Analysis has also reported Luna outperforming Anthropic's Opus 4.8 on this same index, while using roughly a quarter of the estimated cost. That's independent evidence, not marketing, and it cuts in Luna's favor.
Where Luna falls off a cliff
Now the number that actually matters for deciding whether to use Luna: MRCR, short for a long-context recall test that measures a model's ability to accurately find and recall one specific piece of information buried inside a very long document. Think of it as a needle-in-a-haystack test, except the haystack is nearly a million tokens long and the model has to actually pull the right needle out, not just acknowledge the haystack exists.
On MRCR, Luna scores 41.3%. Terra scores 89.6%. Sol scores 91.5%. That's the cliff. It isn't a gentle slope where Luna is "a bit worse" at long context, it's a collapse from near-parity with Terra on general reasoning to being less than half as reliable at the one job a huge context window exists to do. And the score doesn't degrade gradually as documents get longer, either. OpenAI's own reporting has Luna sitting flat at 41.3% whether the document tested is in the 256K-512K token range or the 512K-1M range. It isn't a model that starts strong and fades. It's consistently unreliable at this task, at any length past a certain point.
Here's the sentence worth remembering: a 1,050,000-token context window and a 41.3% long-context recall score are both true of the same model at the same time. Those are two different claims, and it's easy to read the first one and assume it implies the second. It doesn't. Luna can technically accept a huge document. Whether it will actually find the specific fact you need inside that document is a separate question, and for Luna, the honest answer is: often not.
The version trap: read Terminal-Bench numbers carefully
One more thing worth being careful about, because it's a genuine, well-documented trap in how these numbers get reported. On Terminal-Bench 2.1, Luna scores 84.7%, Terra scores 87.4%, and Sol scores 88.8%. Name the version every time you see this benchmark cited, because a separate, harder scoring variant, Terminal-Bench 4.0, produces very different results, Sol scores just 37.3% on that version. A 2.1 score and a 4.0 score are not measuring the same thing and should never be compared or averaged as if they were.
You may also come across a figure of 82.5% for Luna on "Terminal-Bench" with no version stated at all. Treat that number as unverified rather than comparable to the 84.7% figure above, since there's no way to confirm which release it's testing against.
For the complete benchmark set across all three GPT-5.6 tiers, including SWE-bench, GPQA, OSWorld 2.0, and BrowseComp, our GPT-5.6 Sol post covers the full table. We're only pulling the specific numbers here that explain Luna's ceiling.
What GPT-5.6 Luna Actually Costs
This is the table nobody has tabulated properly: every rate across all three GPT-5.6 tiers, in one place, plus the two surcharges that actually change your bill and the price history that explains how these numbers got here.
Tier | Input (standard) | Input (>272K tokens) | Cached input | Cache writes | Output (standard) | Output (>272K tokens) |
|---|---|---|---|---|---|---|
Luna | $0.20 | $0.40 | $0.02 | $0.25 | $1.20 | $1.80 |
Terra | $2.00 | $4.00 | $0.20 | $2.50 | $12.00 | $18.00 |
Sol | $4.00 | $8.00 | $0.40 | $5.00 | $20.00 | $30.00 |
All figures per million tokens.
Tier | Batch/Flex input | Batch/Flex output |
|---|---|---|
Luna | $0.10 | $0.60 |
Terra | $1.00 | $6.00 |
Sol | $2.00 | $10.00 |
Batch and Flex processing runs at roughly half the standard rate on every tier, in exchange for slower, asynchronous turnaround.
The long-context surcharge applies identically across all three tiers, and it's the easiest thing on this whole page to miss until it shows up on an invoice. Any prompt over 272,000 input tokens bills at 2x the input rate and 1.5x the output rate, for the entire request, not just the tokens past the threshold. Cross that line with Luna and your $0.20 input rate becomes $0.40, your $1.20 output rate becomes $1.80. Cache writes are simpler: a flat 1.25x the uncached input rate on every tier, so $0.25 on Luna, $2.50 on Terra, $5.00 on Sol.
How we got here: the price history
July 30, 2026: OpenAI cut Luna's price about 80%, from $1.00 input / $6.00 output to $0.20 / $1.20. Terra was cut about 20% the same day, from $2.50 / $15.00 to $2.00 / $12.00.
August 21, 2026: OpenAI cut Sol's price too, about 20% on input ($5.00 to $4.00) and 33% on output ($30.00 to $20.00), with cached input dropping from $0.50 to $0.40. OpenAI describes this explicitly as promotional pricing, guaranteed for three months, meaning at least through November 21, 2026, not a permanent rate. If you're budgeting a project that runs past that date, build in the possibility that Sol's price reverts toward $5.00 / $30.00.
So, is Luna actually 20 times cheaper than Sol?
At today's prices, yes, on input: $4.00 divided by $0.20 is exactly 20. On output the gap is a bit smaller, about 16.7 times ($20.00 divided by $1.20). You'll see "25x" quoted in some other coverage of this pricing. That figure compares Sol's old launch price ($5.00 / $30.00) to Luna's new, post-cut price ($0.20 / $1.20), mixing two different points in time, which isn't the honest comparison. Compare current prices to current prices and the real gap is 20x on input, roughly 17x on output, and because Sol's rate is itself promotional through at least November 21, that multiple could shrink or grow depending on what OpenAI does after.
The real story of the July 30 cut isn't "the cheap model got cheaper." It's that entire categories of work, high-volume classification pipelines, first-pass drafting, summarizing a support ticket queue, running thousands of calls a day, went from marginal to genuinely economical. At $0.20 per million input tokens, a workload that would have eaten a real budget on last year's flagship pricing can now run at a scale that barely shows up on the invoice. If your cost floor needs to go even lower than that, self-hosted open models are worth a look too, though you're trading OpenAI's infrastructure and support for your own hardware and setup time, our Ollama review covers what that trade actually looks like in practice.
For a deeper look at how per-token price cuts do or don't translate into a lower total bill once you account for retries and heavier usage, our DeepSeek V4 vs Claude Fable 5.1 comparison walks through the actual math on a different pair of models. The same tension applies here: a cheaper token is not automatically a cheaper finished task.
Where Luna Is Strong, Where It Falls Short
Where it's strong:
High-volume, repetitive workloads, classification, tagging, routing, anything where you're making thousands of calls and per-call cost dominates the total bill.
Summarizing short-to-medium documents, well under the 272K long-context surcharge threshold and, more importantly, well under the range where its recall accuracy collapses.
First-pass drafting and lightweight agentic steps, where a human or a stronger model reviews the output afterward anyway.
Latency-sensitive chat and support workflows, it's built to be fast, not just cheap.
General reasoning tasks that don't depend on pulling a specific fact out of a huge document, where it genuinely competes with Terra at a fraction of the price.
Where it falls short:
Any task that depends on accurately finding one specific detail inside a long document or codebase, a 41.3% MRCR score means it will miss the right answer nearly six times out of ten.
Long-context agentic work, multi-file coding tasks, long transcript analysis, anything where the whole point of paying for a huge context window is trusting the model to actually use it.
High-stakes work where a wrong or missed answer is expensive to catch after the fact and nobody is reviewing the output.
Workloads that regularly cross the 272,000-token threshold, at that point you're paying double the input rate and 1.5x the output rate on a model that isn't reliable at that length to begin with, an especially bad combination.
How to Actually Get GPT-5.6 Luna
Through the API, Luna is available now via the model ID gpt-5.6-luna, on the Chat Completions, Responses, and Batch endpoints, using a standard OpenAI API key at developers.openai.com. There's no separate waitlist reported for Luna specifically, being the cheapest and lowest-risk tier in the family, it's the one OpenAI has the least reason to gate behind extra approval.
Outside the raw API, Luna shows up in two places most people will actually touch it: ChatGPT Work subscriptions, where it now consumes fewer credits per use than before the July 30 cut, and the Codex CLI, where OpenAI upgraded the Auto-review feature to run on Luna specifically. If you're on a Work seat or using Codex day to day, you may already be running Luna without having chosen it explicitly.
Who Should Use Luna, Who Should Skip It
Use Luna if: you're running a high-volume pipeline where cost per call genuinely matters, classifying support tickets, tagging content, generating first drafts a human reviews, powering a fast chat interface, or any lightweight agent step where the document or context involved stays well short of the point where recall starts to matter.
Skip Luna if: the job depends on the model finding a specific fact buried in a long document or a large codebase, if you're running an agent that needs to hold and accurately use a genuinely large amount of context over many steps, or if a wrong answer is expensive enough that nobody's reviewing the output before it goes out the door. In any of those cases, the ten-times price premium for Terra is buying real reliability, not just a bigger number on a spec sheet.
That's Luna's own answer. The broader question of which of the three tiers to default to for a given team or workload is bigger than one post can honestly cover here, our GPT-5.6 Terra comparison is where we work through that decision properly.
The Verdict
GPT-5.6 Luna earns its price cut. At $0.20 per million input tokens, it's the cheapest way into the GPT-5.6 family by a wide margin, and on general reasoning it holds up remarkably well against the tier ten times its price. That's not a small claim, and it's backed by both OpenAI's own numbers and an independent index.
But the 41.3% MRCR score is not a footnote, it's the ceiling. Luna is a model you can trust with volume and speed, not with a long document you're relying on it to actually read carefully. If your task needs the second thing, paying more for Terra, or checking whether the specific job actually justifies Sol, is the honest call, not a nice-to-have. Use Luna for what it's built for, high-volume, short-context, cost-sensitive work, and it's arguably the best value in OpenAI's current lineup. Ask it to hold a million tokens in its head and find the one line that matters, and the 20x price gap stops being the interesting number.
Curious to see how it performs?
Try
ChatGPT
Now
GOT ANY QUESTIONS LEFT?
How much does GPT-5.6 Luna cost?
Why did OpenAI cut GPT-5.6 Luna's price by 80%?
Is GPT-5.6 Luna good enough for production work?
What is the cheapest OpenAI model in 2026?


