
Every current Claude model ships with two headline numbers that barely change from one to the next: a 1 million token context window and a knowledge cutoff pushed deep into 2026. Every one of them, except the cheapest one. Claude Haiku 4.5, the $1-per-million-token model sitting at the bottom of Anthropic's lineup, tops out at 200,000 tokens of context, and its reliable knowledge stops in February 2025, about nineteen months behind the rest of the family as of this writing. That gap rarely makes it into the "cheapest Claude model" roundups, and it's exactly the thing you need to know before you build anything durable on top of it.
One housekeeping note before we go further. Anthropic ships new models, pricing changes, and lineup updates at a genuinely fast pace, four current-generation flagship models in roughly a year at last count. Everything below reflects Anthropic's own documentation as of September 12, 2026. Confirm the current model lineup, pricing, and specs directly on platform.claude.com before you commit a production decision to any specific number here.
So what is Claude Haiku 4.5, in one line? It's Anthropic's fastest, cheapest current model, priced at $1 per million input tokens and $5 per million output tokens, built for high-volume, latency-sensitive work rather than heavy reasoning. Is it worth it? For the jobs it was actually built for, yes, easily. For anything that needs to hold more than a book's worth of context in one request, needs knowledge newer than early 2025, or needs to keep running untouched for the next two years without a migration, the honest answer gets more complicated. That's the rest of this review.
Claude Haiku 4.5 at a Glance
Spec | Detail |
|---|---|
Maker | Anthropic |
Release date | October 15, 2025 |
Model ID | claude-haiku-4-5-20251001 (alias: claude-haiku-4-5) |
Context window | 200,000 tokens (roughly 150,000 words) |
Max output | 64,000 tokens |
Reliable knowledge cutoff | February 2025 |
Training data cutoff | July 2025 |
Input price | $1 per million tokens |
Output price | $5 per million tokens |
Thinking mode | Extended (manual budget), not Adaptive |
Status | Active; retirement not sooner than October 15, 2026 |
Availability | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS |
That's the whole shape of the model in one table, pulled straight from Anthropic's own model overview. Two rows are worth sitting with before you read on: the context window and the retirement date are both the smallest numbers anywhere in Anthropic's current lineup, and neither one shows up on the price tag.
What Claude Haiku 4.5 Is, and Why It Exists
Every model family needs a workhorse tier, the option built for the request that's too routine to justify flagship pricing, sent by the thousand rather than composed carefully one at a time. That's the role Haiku fills in Anthropic's lineup, and Haiku 4.5 is the current version of it.
It isn't a new, untested entrant sitting at the bottom of the family either. It's the tier's proven incumbent. When Anthropic retired Claude Haiku 3.5 on February 19, 2026, and Claude Haiku 3 on April 20, 2026, its own deprecation notices pointed every migrating developer toward the same replacement: claude-haiku-4-5-20251001. If you were running either of those older, cheaper models in production, Haiku 4.5 is where Anthropic already moved you, not a sidelined option you'd be taking a risk on today.
Above Haiku sits Claude Sonnet 5, which Anthropic itself describes as "the best combination of speed and intelligence" in the current lineup, and we've covered that model's own tier in a separate Claude Sonnet 5 review. Above that again sit Claude Opus 5 and Claude Fable 5.1, built for heavier agentic and reasoning work. Rather than repeat a full family-wide pricing and specs table here, that complete four-model comparison lives in our Claude Opus 5 review, so we'll stay focused on what actually makes Haiku 4.5 different.
What's Actually New
Speed is the headline, and it holds up. Anthropic's own model comparison table describes Haiku 4.5 as "the fastest model with near-frontier intelligence" in its current lineup, and in practice it responds noticeably quicker than Sonnet 5 or Opus 5 on the same prompt. That matters more than it sounds like for anything with a human or another system waiting on the other end: a live chat widget, a voice agent, an autocomplete suggestion, a routing step inside a bigger pipeline.
Here's a technical difference that actually changes how you'd build with it, though. Every other model in Anthropic's current lineup, Sonnet 5, Opus 5, and Fable 5.1, uses what Anthropic calls Adaptive thinking: you set an effort level and the model decides for itself how many reasoning tokens that's worth. Haiku 4.5 doesn't get that dial at all, the effort parameter simply isn't supported on it. It's stuck on Extended thinking, the older, manual mode where you set thinking.type: "enabled" and hand the model a fixed budget_tokens number yourself. Think of Adaptive thinking as cruise control: tell it roughly how fast, and the car handles the rest. Extended thinking is closer to a manual gearbox. It works fine, but you're the one doing the extra work on every request.
Availability is the other genuine strength. Haiku 4.5 is offered on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS, the broadest platform spread of any current Claude model, which matters if your infrastructure is already committed to one of those clouds specifically.
What the Benchmarks Actually Say
This is the section a vendor launch chart wants you to skim past, so let's not. Anthropic's own numbers and independent testing tell two related but different stories here, and they need to stay clearly separated rather than blended into one impressive-sounding paragraph.
What Anthropic reports
On its own launch announcement, Anthropic reports Haiku 4.5 scoring 73.3% on SWE-bench Verified, a widely used real-world coding benchmark. The company is specific about the setup behind that number: "a simple scaffold with two tools, bash and file editing via string replacements," averaged over 50 trials, with a 128K thinking budget, no test-time compute, default sampling parameters, and the full 500-problem dataset. That's a real, disclosed methodology, worth noting since plenty of benchmark numbers in this industry arrive without one.
Anthropic's own framing of that score is that Haiku 4.5 delivers coding performance "similar" to Claude Sonnet 4, a model that was itself state of the art five months before Haiku 4.5 shipped, but "at one-third the cost and more than twice the speed." That's a comparison against the previous Sonnet generation, not the current Sonnet 5, and it's Anthropic's own comparison, worth stating plainly since it's the vendor's chart, not an outside lab's. Separately, Augment, a coding-tools company and named partner rather than an independent benchmark lab, reported Haiku 4.5 reaching about 90% of Sonnet 4.5's performance on its own internal agentic coding evaluation, a useful data point, but again, not an independent, published benchmark result.
What independent testing shows
Artificial Analysis, a third-party benchmark aggregator with no stake in Anthropic's own marketing, puts Haiku 4.5 (non-reasoning) at an estimated 15 on its Intelligence Index, against Claude Sonnet 4.5 (non-reasoning) at an estimated 19. In reasoning mode, the gap is 30 versus 36. That's a real, independently measured difference, and it says something Anthropic's own coding-specific numbers don't: Haiku trails Sonnet by a meaningful margin on general reasoning, even while holding up well on coding tasks specifically.
The gap nobody's filled in yet
Here's the honest limitation running through all of the above: every independent comparison of Haiku 4.5 available right now, including Artificial Analysis' numbers, pits it against Claude Sonnet 4.5, the previous Sonnet generation, not Claude Sonnet 5, the model actually sitting above Haiku in today's lineup. Sonnet 5 shipped after Haiku 4.5's original benchmarks were run, with a different context window, price, and knowledge cutoff attached to it. On paper, per Anthropic's own model overview and pricing page, Sonnet 5 costs twice as much per token ($2/$10 versus $1/$5), offers five times the context window (1 million versus 200,000 tokens), and carries roughly a year of extra knowledge (January 2026 versus February 2025). But an actual independent benchmark score pitting Haiku 4.5 against the current Sonnet 5, run by a third party, doesn't appear to exist yet as of this writing. That's worth knowing plainly rather than papering over with an older comparison and presenting it as current.
What Claude Haiku 4.5 Actually Costs
The headline price is straightforward: $1 per million input tokens, $5 per million output tokens, on the standard Claude API. A few modifiers change the real bill, though, and they're worth knowing before you estimate a budget.
Prompt caching cuts costs on repeated context. A 5-minute cache write costs 1.25 times the base input price ($1.25 per million on Haiku), a 1-hour cache write costs twice the base price ($2 per million), and a cache read, once that content is actually reused, costs just 10% of the base input price, $0.10 per million tokens. That 10% rate is the standard multiplier across almost the entire Claude lineup; the much cheaper 2.5% cache-read rate is specific to Claude Fable 5.1 and Claude Mythos 5.1 and doesn't apply to Haiku, worth flagging since the two numbers get confused easily.
The Batch API, for anything that doesn't need a synchronous response, cuts both input and output prices in half: $0.50 per million input, $2.50 per million output. Anthropic's own pricing documentation runs a worked example that shows what this looks like at scale: about 3,700 tokens per conversation across 10,000 customer support tickets, on Haiku 4.5 at standard rates, comes out to roughly $37 total. That's the kind of workload Haiku is actually built for.
Flip the usual "is the flagship worth the extra cost" question around here, because Haiku is already the cheap option. The real question most readers have is the opposite one: when is it worth paying Sonnet 5's $2/$10 instead of Haiku's $1/$5? A rule that holds up in practice: any single request needing more than 200,000 tokens of input, more than 64,000 tokens of output, or knowledge newer than February 2025 simply can't run on Haiku at all, so the "is it worth double" question doesn't come up, there's no choice to make. Below those ceilings, on work Haiku can technically handle, Sonnet 5's Adaptive thinking and finer reasoning control buy you meaningfully better output at double the token price. For the complete four-model pricing table, Haiku 4.5 alongside Sonnet 5, Opus 5, and Fable 5.1, see our Claude Opus 5 review, which owns that full comparison.
Where It's Strong, Where It Falls Short
Where it's strong:
The cheapest and fastest model in Anthropic's current lineup, at $1 per million input tokens and $5 per million output tokens
Near-frontier coding performance for a fraction of Sonnet-tier cost, 73.3% on SWE-bench Verified by Anthropic's own disclosed methodology
A proven incumbent, not an unproven newcomer, it is the exact model Anthropic pointed retiring Haiku 3 and Haiku 3.5 customers toward
The broadest platform reach of any current Claude model: Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS all offer it
Batch and prompt-cache discounts stack on an already-low base price, pushing high-volume workloads even lower
Where it falls short:
A 200,000-token context window, the smallest of any current Claude, against 1 million tokens on Sonnet 5, Opus 5, and Fable 5.1
A 64,000-token max output limit, half the 128,000-token ceiling every other current model offers
A reliable knowledge cutoff of February 2025, roughly nineteen months stale as of this writing and the oldest in the current lineup
Extended thinking only, with no Adaptive thinking or effort parameter, meaning you manage reasoning budgets by hand instead of letting the model decide
By far the nearest retirement floor of any active Claude model. To be precise about what that does and does not mean: Haiku 4.5 is still listed Active, not Deprecated, on Anthropic's model deprecations page, and no retirement notice has actually been issued. Anthropic commits to at least 60 days notice before retiring any publicly released model. But its stated commitment is "not sooner than October 15, 2026," roughly five weeks from this article's publish date, against 2027 dates for every other model in the current lineup. That is a floor, not an announced end date, but it is the nearest one by a wide margin, and it is worth knowing before you build something meant to last on the cheapest Claude.
How to Actually Get Claude Haiku 4.5
Through the API, Haiku 4.5 is available now under the model ID claude-haiku-4-5-20251001, or the alias claude-haiku-4-5, using a standard Claude API key. It is also available through Amazon Bedrock (anthropic.claude-haiku-4-5), Google Cloud (claude-haiku-4-5@20251001), Microsoft Foundry, and Claude Platform on AWS, so if your infrastructure already sits on one of those clouds, you are not locked into going through Anthropic directly to use it.
One thing worth flagging honestly rather than guessing at: whether Haiku 4.5 shows up as a selectable model inside the consumer claude.ai chat interface for Pro or free users is not something Anthropic's developer documentation actually addresses, that is a product decision for the consumer app, separate from API availability. If you are evaluating it for a chat use case rather than building against the API directly, check the model picker inside claude.ai itself rather than assuming based on API access.
Who Should Use It, Who Should Skip It
Use Haiku 4.5 if you are running a high-volume pipeline where cost per call genuinely adds up: classifying support tickets, tagging content, routing requests, powering a fast chat widget, or generating first-pass drafts that a human or a stronger model reviews afterward. It is also a solid pick anywhere response latency matters more than squeezing out the last point of reasoning quality.
Skip it if the job depends on holding a long document or a large codebase in context at once, needs knowledge from anywhere in the last year and a half without heavy retrieval-augmented grounding, or if you are building infrastructure meant to run untouched for the next couple of years without a migration plan already in mind. In any of those cases, the extra cost of Sonnet 5 is buying real capability, not just a bigger number on a spec sheet.
If you are weighing this against OpenAI's equivalent budget tier rather than staying inside the Claude family, we have reviewed the closest comparison separately: GPT-5.6 Luna, OpenAI's own cheapest current model. Both chase the same "good enough, priced low" spot in their respective lineups, and seeing them side by side is a genuinely useful sanity check before picking either one by default.
The Verdict
Claude Haiku 4.5 earns its price tag. At $1 per million input tokens and $5 per million output tokens, it is the cheapest way into the current Claude family, and on the coding benchmarks Anthropic actually discloses its methodology for, it holds up well against models that cost several times more.
But the trade-offs are not cosmetic. A 200,000-token context window, a February 2025 knowledge cutoff, and the nearest retirement floor of any active Claude model are all real limitations, not footnotes, and none of them show up in the headline price. Use Haiku 4.5 for what it is actually built for, high-volume, cost-sensitive, latency-sensitive work well within its context ceiling, and it is arguably the best value in Anthropic's current lineup. Ask it to hold a long document in memory, work with knowledge from the last year, or anchor something you do not plan to touch again for years, and Sonnet 5's extra cost stops being optional.
Curious to see how it performs?
Try
Claude
Now
GOT ANY QUESTIONS LEFT?
What is Claude Haiku 4.5?
How much does Claude Haiku 4.5 cost?
What is Claude Haiku 4.5's context window?
Is Claude Haiku 4.5 being discontinued?


