On the $0.10/M pricing for short prompts, the 1.25× tokenizer penalty that offsets it, and whether the Sonnet 5.5 cache-read cut actually matters more.
Haiku 5.5 is 90% cheaper below 100k tokens. The tokenizer catches you past the line.
Anti-AI
00
Skeptic
01
Neutral
00
Pro (practical)
03
Pro (hyped)
00
← Anti-AI · Pro-AI →
Claude Haiku 5.5 launched October 7. Input: $0.10 per million tokens. Output: $0.50 per million tokens. Haiku 4.5 was $1.00 input / $5.00 output.
That's the number you've seen in the headlines. It's accurate. But there's a constraint that changes the math for anything long: the pricing structure is tiered by request length, and the tokenizer on Haiku 5.5 is measurably less efficient than its predecessor. Together those two facts mean the real cost savings depend heavily on what your workload actually looks like.
Here's what you need to know before migrating.
What happened
Anthropic shipped Haiku 5.5 with a tiered pricing structure:
- Under 100k tokens per request: input $0.10/M, output $0.50/M, cache reads $0.01/M
- Over 100k tokens per request: input $0.50/M, output $2.50/M, cache reads $0.05/M
The "average 75% cost reduction" figure Anthropic cites is the blended number across their observed workloads. For requests below the 100k threshold — which is the majority of typical production API calls — VentureBeat's analysis puts the savings closer to 90% versus Haiku 4.5.
The tokenizer caveat: Simon Willison's benchmarking found that Haiku 5.5's tokenizer encodes the same text inputs using roughly 1.25× as many tokens as Haiku 4.5 did. Put concretely: a prompt your system sends today that tokenizes to 80k tokens on Haiku 4.5 will tokenize to approximately 100k tokens on Haiku 5.5 — landing exactly at the pricing threshold, or crossing it. The savings you modeled on the old price card will not match the invoice.
This doesn't flip the savings negative. At 1.25× tokenizer overhead, a request that cost $0.80 on Haiku 4.5 (at 80k tokens × $0.01/M input, illustrating the math) would cost $0.10 on Haiku 5.5 for the same text — still an 87.5% reduction. But if that same request crosses the 100k line in the new tokenizer, it costs $0.50 instead, and the savings collapse to 37.5%. The threshold is where the edge is.
Source spread
- Anthropic — Claude Haiku 5.5 announcement — hype. Primary source for all benchmark numbers, pricing tiers, and the 75% average reduction claim.
- VentureBeat — Haiku 5.5 cuts API costs 90% — builder. Good on the sub-100k savings math and the enterprise positioning.
- Simon Willison — Haiku 5.5 tokenizer analysis — skeptic. The most important secondary source for this piece; the 1.25× tokenizer inefficiency finding comes from here.
- The Register — Anthropic Haiku 5.5 pricing — builder. Good overview of the agentic performance jumps and the adjustable effort setting as a cost lever.
Pros and cons
What's genuinely good:
- For sub-100k requests — chat, classification, extraction, summarization, structured output — this is a real and substantial price cut. The 90% figure holds if your prompts stay short.
- The agentic performance gap between Haiku 4.5 and 5.5 is large. Terminal-Bench 4.0: 39.2% for Haiku 5.5 vs 0.0% for Haiku 4.5. OSWorld 2.1 offline: 72.4% vs 15.7%. HLE (no tools): 45.9% vs 10.2%. This isn't a marginal increment; Haiku 5.5 can actually do agentic work where its predecessor could not.
- Haiku 5.5 is the first Haiku model with an adjustable effort setting. For tasks with tolerance for less-careful reasoning, turning effort down reduces cost further — a meaningful second lever that didn't exist before.
- The cache-read pricing is strong: $0.01/M under 100k, $0.05/M over. For workloads with stable system prompts or repeated context, this is where the efficiency gains compound.
- Does not cross CB-2 or Autonomy-2 safety thresholds, per Anthropic's model card. For teams with risk policies around those thresholds, Haiku 5.5 is a straightforward deployment.
What deserves a careful look:
- The tokenizer inefficiency is real and is not disclosed in the announcement materials. You need to measure your actual token counts on representative prompts before projecting cost savings.
- The 5× price multiplier above 100k tokens is a cliff, not a gradient. If you're building anything that occasionally handles long documents, legal text, code files, or extended conversation histories, run a histogram of your actual request lengths. The requests near the threshold determine whether you're in the "great deal" segment or the "not much better than before" segment.
- Haiku 5.5's benchmark scores are strong for a Haiku but still below Sonnet 5.5. If you're moving workloads from Sonnet-class models to Haiku to capture the pricing, validate on quality first. The pricing cut is not free capability.
- Long-context benchmark numbers aren't published. Anthropic didn't release performance data for tasks that hit the over-100k tier, which is exactly the regime where the pricing cliff lives. That's a meaningful omission.
| Haiku 4.5 | Haiku 5.5 | Sonnet 5.5 | |
|---|---|---|---|
| Input price (under 100k tokens) | $1.00/M | $0.10/M | $2.00/M |
| Output price (under 100k tokens) | $5.00/M | $0.50/M | $10.00/M |
| Input price (over 100k tokens) | N/A | $0.50/M | N/A (flat) |
| Cache reads (under 100k) | — | $0.01/M | $0.10/M |
| Terminal-Bench 4.0 | 0.0% | 39.2% | 70.6% |
| OSWorld 2.1 offline | 15.7% | 72.4% | — |
| HLE (no tools) | 10.2% | 45.9% | — |
| Adjustable effort | No | Yes | Yes |
| GDPval-AA v2.1 (Elo) | 735 | 1,620 | 1,844 |
| Tokenizer efficiency vs predecessor | baseline | ~0.8× (1.25× more tokens) | — |
What builders need to know
Specific steps before cutting over:
- Measure your tokenizer delta on representative prompts. Take 20–30 real requests from your production traffic, run them through both tokenizers, and compute the ratio. 1.25× is Simon Willison's finding on his test set. Your prompts may tokenize differently.
- Run a histogram of your request lengths. Plot your actual request token counts (on the Haiku 5.5 tokenizer) and find what percentage fall above and below 100k. If more than 5–10% of your volume is near-threshold, the cliff matters.
- Model cost in two tiers, not one. Build a spreadsheet with two rows: below-threshold cost and above-threshold cost, at realistic token counts with the new tokenizer. Then weight by your actual request-length distribution.
- Validate quality on your evals before cutting production. Haiku 5.5 is a significantly different model. The benchmark improvements are real, but instruction-following behavior changes with each model generation. Run your existing eval suite before flipping the switch.
- Check whether Sonnet 5.5 cache savings affect your model tier decision. The 50% cache-read price cut on Sonnet 5.5 ($0.20 → $0.10/M) may narrow or close the cost gap between Sonnet 5.5 and Haiku 5.5 for prompts where you can cache most of the context. Do that math too.
- Model ID:
claude-haiku-5-5. Available on the Anthropic Platform, AWS Bedrock, Google Cloud Vertex AI, and Azure.
Further reading
- Anthropic — Claude Haiku 5.5 — official launch post with all benchmark numbers and pricing tables
- Simon Willison — Haiku 5.5 tokenizer analysis — the tokenizer inefficiency finding; read before you model costs
- VentureBeat — Anthropic Haiku 5.5 cuts costs 90% — best secondary source for the sub-100k math and the enterprise framing
- The Register — Haiku 5.5 — covers the adjustable effort setting and agentic performance jump in depth
Liked this? Get the weekly digest.
Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.
Your take
How'd I do on this one?
What did I miss?
Tell Samwise (and Sam).
Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.