AI Model Reviews

GPT-5.6 Terra vs Sol vs Luna: Which Tier Should You Pay For?

GPT-5.6 Terra vs Sol vs Luna featured image comparing OpenAI’s flagship, balanced, and affordable AI model tiers, with Terra highlighted as the best balance

Ask ChatGPT which GPT-5.6 model to reach for and the reflex answer is usually Sol, because Sol is the flagship and flagships are what you are supposed to want. That reflex costs money for no good reason. GPT-5.6 Terra, the tier sitting quietly in the middle of OpenAI's July 2026 lineup, scores within 1.4 points of Sol on Terminal-Bench 2.1 and within 1.9 points on long-context recall, at close to half Sol's price. Terra also barely shows up in its own search results, mostly buried under API pricing pages on Azure, Amazon Bedrock, and OpenRouter. Nobody is writing about whether you should actually use it. This post is.

One note before the numbers: OpenAI has already touched this exact lineup three times since launch, cutting Terra's and Luna's prices in a single day, discounting Sol separately a few weeks later, and then shipping a new flagship, GPT-6 Astra, on top of all three. Treat everything below as accurate as of September 8, 2026, and check current pricing on OpenAI's own docs before you commit real budget to any of it.

This is not a straight review of one model, and it is not another retyped launch chart either. It is the actual decision: when GPT-5.6 Sol, Terra, and Luna are all sitting in the same dropdown, which one should you be paying for, and when does each one genuinely earn its keep.

GPT-5.6 Terra at a Glance

Spec

Detail

Maker

OpenAI

Released

July 9, 2026

Model ID

gpt-5.6-terra

Position in lineup

Mid-tier, between Sol (flagship) and Luna (fast/affordable)

Context window

1,050,000 tokens

Max output

128,000 tokens

Input price

$2 per million tokens

Output price

$12 per million tokens

Input types

Text and image

Output types

Text only

Availability

OpenAI API, Amazon Bedrock, Microsoft Foundry (Azure)

What GPT-5.6 Terra Actually Is

GPT-5.6 launched July 9, 2026 as three separate models sharing one generation number: Sol, Terra, and Luna. OpenAI's own framing is that the number identifies the generation, while the names identify durable capability tiers that can each move on their own release schedule going forward, rather than one flagship model getting swapped out every few months. Terra is the middle tier, not the biggest, not the cheapest, built for what OpenAI calls everyday production work: code generation, content workflows, structured data extraction, and general-purpose agentic tasks that do not need Sol's absolute ceiling.

That framing matters because it changes how you should think about picking a tier. This is not "good, better, best" with Terra as the compromise pick. It is three models built for different jobs, and for a lot of real work, the middle one is not a compromise, it is the correct tool. That is the entire argument this post is here to make.

What Terra Actually Supports

Terra ships with the same core toolset most builders actually reach for: streaming, structured outputs, function calling, file search, image input, web search, and prompt caching, all through the Responses, Chat Completions, and Batch endpoints. What it does not do is just as telling. No Realtime, Assistants, fine-tuning, embeddings, image generation, video, speech, or the legacy Completions endpoint. Terra is a text-and-image-in, text-out reasoning and agentic model, not a general-purpose multimodal hub. If your workload needs any of those missing endpoints, you are reaching for a different model in OpenAI's catalog regardless of which GPT-5.6 tier you pick.

Third-party analysis of the July launch, Vellum's own writeup among others, describes Terra as delivering GPT-5.5-class performance at under half Sol's price on most measures, and notes it surpasses GPT-5.5's own peak scores on OSWorld and BrowseComp specifically, not just matching them. Worth flagging plainly: that framing traces back to OpenAI's own launch materials, not an independent benchmark run, so treat it as directional rather than settled fact. The numbers below are where the real comparison lives.

GPT-5.6 Terra vs Sol: How Close Is the Real Gap?

Here is where a lot of coverage of this launch gets sloppy, and it is worth being careful, because two versions of the same benchmark are circulating right now with wildly different scales. Terminal-Bench 2.1 puts Sol at 88.8%. Terminal-Bench 4.0, a separate, much harder benchmark released later, puts Sol at 37.3%. Same benchmark family, different task sets, not comparable numbers. Every Terminal-Bench figure in this post is 2.1, stated plainly, so it never gets mixed with the other scale.

On that basis, OpenAI's own published Terminal-Bench 2.1 results put Terra at 87.4% against Sol's 88.8%, a 1.4 point gap. On MRCR v2, an 8-needle long-context recall test run in the 256K to 512K token band, OpenAI's own numbers put Terra at 89.6% against Sol's 91.5%, a 1.9 point gap. Those are both OpenAI's own vendor-reported figures, not independent, but they line up with an outside check: Artificial Analysis, an independent benchmark aggregator, puts Terra at 77.4 and Sol at 80 on its own Coding Agent Index, a 2.6 point gap in the same direction, from a source with nothing to gain by flattering OpenAI's mid-tier model.

Put a dollar figure on that gap. A request with 1 million input tokens and 100,000 output tokens costs about $6 on Sol and about $3.20 on Terra. Terra runs roughly 53% of Sol's cost for a task where the two models land within two points of each other on every measure above. That is the entire case for Terra in one sentence: you are paying double for a gap you can barely find on the scoreboard. For the complete picture, including the benchmarks this post deliberately leaves out, see GPT-5.6 Sol's full benchmark results.

Sol does still win outright on one measure worth naming honestly. Agents' Last Exam, OpenAI's hardest agentic reasoning test, has Sol scoring 53.6 against Terra's 50.4, a real 3.2 point gap, wider than the coding and recall gaps above. If your work looks like Agents' Last Exam, long, hard, multi-step reasoning where a wrong turn compounds, that gap is the one that should actually change your decision, not the smaller ones.

GPT-5.6 Terra vs Luna: Where the Cheap Tier Falls Apart

This comparison matters more than most people expect, because on paper Luna looks like it should win on value. Luna costs $0.20 input and $1.20 output, a tenth of Terra's price. And on OpenAI's own Agents' Last Exam, Luna scores 50.3 against Terra's 50.4, a gap of one tenth of a point. Statistically, on that one measure, they are the same model.

Then look at MRCR, the long-context recall test. Terra holds 89.6%. Luna drops to 41.3%. That is not a gap, it is a cliff, more than 48 points, on a test measuring whether the model can actually find and use something you told it earlier in a long conversation or document. Independent scoring backs the same shape: Artificial Analysis' Coding Agent Index has Terra at 77.4 and Luna at 74.6, a modest 2.8 point difference, close to the Agents' Last Exam parity, not the MRCR cliff.

Put those two facts together and you get the single most useful rule in this whole comparison. On short, contained reasoning tasks, Luna genuinely matches Terra. The moment a task needs the model to hold and use something from earlier in a long conversation or a long document, Luna falls off a cliff and Terra does not. If you are running short customer support replies, quick classification, or simple extraction, Luna is not a compromise, it is the correct, cheaper choice. If you are running anything that spans a long document, a long chat history, or a large codebase, the ten-cent price gap stops mattering the moment Luna starts missing context it should have caught.

What Each GPT-5.6 Tier Costs

The three headline rates, current as of this writing: Sol runs $4 input and $20 output per million tokens, a promotional rate in effect since an August 21, 2026 cut from $5/$30, guaranteed to hold at least through November 21, 2026. Terra runs $2 input and $12 output, down about 20% from $2.50/$15 after a July 30, 2026 cut. Luna runs $0.20 input and $1.20 output, down roughly 80% from $1/$6, cut the same day as Terra's.

One billing rule applies to all three tiers, and it is easy to get burned by it. Cross 272,000 input tokens in a single request, and the entire request, not just the tokens past that line, gets re-rated at 2x the input rate and 1.5x the output rate. A 300,000-token prompt on Terra does not pay the extra 28,000 tokens at double price, it pays for all 300,000 tokens at double price. Budget around that threshold deliberately if your workload runs long prompts, on any of the three tiers, not just Terra. For the complete pricing picture across all three tiers, including cached-token rates, batch pricing, and the long-context surcharge math worked out plan by plan, see the full GPT-5.6 pricing breakdown.

Where Terra Is Strong, Where It Falls Short

Where it's strong:

  • Within 1.4 to 1.9 points of Sol on Terminal-Bench 2.1 and MRCR, at roughly half Sol's blended cost.

  • Independently confirmed to trail Sol by a similarly small margin, 2.6 points, on Artificial Analysis' own Coding Agent Index, not just on OpenAI's own numbers.

  • Full agentic tool support: streaming, structured outputs, function calling, file search, web search, and prompt caching, the complete stack, not a stripped-down cheap tier.

  • A genuinely large 922,000-token max input, enough for most real codebases and long documents in a single request.

  • On short reasoning tasks it is barely distinguishable from Luna at a tenth of the price, but holds up on long-context work where Luna does not.

Where it falls short:

  • Still measurably behind Sol on the hardest agentic reasoning work, a real 3.2 point gap on Agents' Last Exam that matters more the harder the task gets.

  • Text and image input, text-only output, no realtime, fine-tuning, embeddings, image generation, or speech endpoints, a narrower tool than OpenAI's broader catalog suggests.

  • The 272,000-token billing threshold re-rates the entire request, easy to get an unpleasant surprise on a long-prompt workload.

  • Not directly selectable in the standard ChatGPT consumer app model picker, so most non-developers cannot just click their way to using it.

  • Too recent, launched July 2026, for the kind of multi-month independent production track record older models already have.

How to Actually Get GPT-5.6 Terra

Terra is not something most consumer ChatGPT users can select from a dropdown. It is not one of the models in the standard ChatGPT model picker on Plus, Pro, or Team plans. It does show up inside ChatGPT Work and inside Codex, OpenAI's coding agent product, including on Free and Go tiers there, which is currently the most generous path to using Terra without touching the API directly.

For everyone else, Terra is an API model, gpt-5.6-terra, through OpenAI's own Responses, Chat Completions, and Batch endpoints, through Amazon Bedrock on either the Bedrock Converse API or the OpenAI-compatible endpoints, or through Microsoft Foundry on Azure. That is three separate clouds carrying the same model, which is part of why Terra's own name is so crowded with API documentation pages and so light on plain-English explanations of when to actually use it.

Worth naming honestly: GPT-6 Astra, OpenAI's new flagship, launched September 3, 2026 at $10 input and $50 output, roughly five times Terra's input price. Astra is not part of this comparison in the tier-for-tier sense, it sits above Sol, not between Terra and Luna, and for most everyday work the more useful mental model is routing rather than defaulting to whichever tier launched most recently: Luna for easy, high-volume turns, Terra for the bulk of everyday work, Sol for the genuinely hard steps, Astra reserved for the slice of tasks Sol demonstrably cannot handle. If you are weighing switching platforms entirely rather than picking a tier inside one, our ChatGPT vs Claude comparison and our look at whether AI is actually getting cheaper both cover that broader question.

Which GPT-5.6 Tier Should You Use? The Real Decision Framework

This is the actual question this post exists to answer, so here is a direct answer rather than a hedge.

Your situation

Use this tier

Why

Short replies, classification, simple extraction, high volume

Luna

Matches Terra on short reasoning tasks at a tenth of the price; the MRCR gap only bites on long-context work

Everyday coding, content generation, agentic tasks, most production work

Terra

Within 2 points of Sol on the benchmarks that matter for this, at roughly half the cost

Long documents, long chat history, large codebases in a single request

Terra, not Luna

Luna's MRCR score collapses to 41.3% on exactly this kind of task; Terra holds 89.6%

Genuinely hard, multi-step agentic reasoning where a wrong step compounds

Sol

The 3.2 point Agents' Last Exam gap over Terra is real and widens on the hardest tasks

Security research, frontier-level reasoning tasks Sol cannot handle

GPT-6 Astra

Sits above Sol entirely; reserve it for the slice of work Sol demonstrably fails, not as a default

The pattern underneath that table is simple: default to Terra, not Sol. Drop to Luna only when the task is genuinely short and self-contained. Escalate to Sol only when you have a specific reason tied to Agents' Last Exam-style hard reasoning, not just a vague sense that the flagship should be safer. Most people default upward out of habit, not because the numbers ask them to.

Who Should Use Terra, Who Should Skip It

Use Terra by default if you are running everyday production work, coding tasks, content generation, structured extraction, or general agentic workflows, and you have not specifically confirmed that Sol's extra points on any given benchmark matter to your outcome. Also use it over Luna the moment your task involves a long document, a long conversation history, or a large codebase, since that is exactly where Luna's MRCR score falls apart and Terra's does not.

Skip Terra and go to Sol if your workload resembles Agents' Last Exam, long, hard, multi-step reasoning where a wrong turn early on compounds into a bigger failure later, or if you are running tasks at the edge of what any GPT-5.6 tier can do and the 3.2 point gap is the difference between a task finishing correctly and not. Skip Terra and drop to Luna if your task is short, contained, and repeated at real volume, first-pass classification, short replies, simple extraction, where the ten-cent-per-million price gap actually adds up and the MRCR weakness never gets triggered because nothing you are doing needs long-context recall.

Verdict: Pay for Terra, Not Sol, Most of the Time

The honest read on GPT-5.6's three tiers is that OpenAI built a genuinely good middle option and then let its own flagship's name do all the marketing work, the same pattern now playing out with Astra sitting on top of all three. Terra is not a compromise pick. It lands within 1.4 to 1.9 points of Sol on the benchmarks that cover most real work, an independent aggregator confirms a similarly small gap rather than a vendor-flattering one, and it does all of that at roughly half Sol's blended cost.

Sol still earns its price on genuinely hard, multi-step agentic reasoning, where its lead over Terra is real and widens rather than shrinks. Luna earns its place on short, high-volume, contained tasks, but its long-context collapse means it is not a safe universal substitute for Terra the way its price tag alone might suggest. If you are defaulting to Sol out of habit rather than a specific reason tied to your hardest tasks, that habit is costing you money for a gap you would struggle to notice in practice. Start with Terra. Move up only when you can point to the specific number that says you need to.

Curious to see how it performs?

Try

Claude

Now

GOT ANY QUESTIONS LEFT?

Which GPT-5.6 tier should I use?

What is the difference between GPT-5.6 Sol, Terra and Luna?

Do Sol, Terra and Luna share the same context window and knowledge cutoff?

Is GPT-5.6 Terra good enough to replace Sol?