Vol. 1 · Edition 041Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

$0.10
per million input tokens
Haiku 5.5 — under 100k token threshold
Model Launch
By Sam Taylor with Samwise

On the $0.10/M pricing for short prompts, the 1.25× tokenizer penalty that offsets it, and whether the Sonnet 5.5 cache-read cut actually matters more.

Haiku 5.5 is 90% cheaper below 100k tokens. The tokenizer catches you past the line.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

00

Pro (practical)

03

Pro (hyped)

00

← Anti-AI · Pro-AI →

Claude Haiku 5.5 launched October 7. Input: $0.10 per million tokens. Output: $0.50 per million tokens. Haiku 4.5 was $1.00 input / $5.00 output.

That's the number you've seen in the headlines. It's accurate. But there's a constraint that changes the math for anything long: the pricing structure is tiered by request length, and the tokenizer on Haiku 5.5 is measurably less efficient than its predecessor. Together those two facts mean the real cost savings depend heavily on what your workload actually looks like.

Here's what you need to know before migrating.

What happened

Anthropic shipped Haiku 5.5 with a tiered pricing structure:

  • Under 100k tokens per request: input $0.10/M, output $0.50/M, cache reads $0.01/M
  • Over 100k tokens per request: input $0.50/M, output $2.50/M, cache reads $0.05/M

The "average 75% cost reduction" figure Anthropic cites is the blended number across their observed workloads. For requests below the 100k threshold — which is the majority of typical production API calls — VentureBeat's analysis puts the savings closer to 90% versus Haiku 4.5.

5×
Price jump from below to above the 100k token threshold — input goes from $0.10/M to $0.50/M, output from $0.50/M to $2.50/M

→ Source: Anthropic

The tokenizer caveat: Simon Willison's benchmarking found that Haiku 5.5's tokenizer encodes the same text inputs using roughly 1.25× as many tokens as Haiku 4.5 did. Put concretely: a prompt your system sends today that tokenizes to 80k tokens on Haiku 4.5 will tokenize to approximately 100k tokens on Haiku 5.5 — landing exactly at the pricing threshold, or crossing it. The savings you modeled on the old price card will not match the invoice.

This doesn't flip the savings negative. At 1.25× tokenizer overhead, a request that cost $0.80 on Haiku 4.5 (at 80k tokens × $0.01/M input, illustrating the math) would cost $0.10 on Haiku 5.5 for the same text — still an 87.5% reduction. But if that same request crosses the 100k line in the new tokenizer, it costs $0.50 instead, and the savings collapse to 37.5%. The threshold is where the edge is.

Source spread

Pros and cons

What's genuinely good:

  • For sub-100k requests — chat, classification, extraction, summarization, structured output — this is a real and substantial price cut. The 90% figure holds if your prompts stay short.
  • The agentic performance gap between Haiku 4.5 and 5.5 is large. Terminal-Bench 4.0: 39.2% for Haiku 5.5 vs 0.0% for Haiku 4.5. OSWorld 2.1 offline: 72.4% vs 15.7%. HLE (no tools): 45.9% vs 10.2%. This isn't a marginal increment; Haiku 5.5 can actually do agentic work where its predecessor could not.
  • Haiku 5.5 is the first Haiku model with an adjustable effort setting. For tasks with tolerance for less-careful reasoning, turning effort down reduces cost further — a meaningful second lever that didn't exist before.
  • The cache-read pricing is strong: $0.01/M under 100k, $0.05/M over. For workloads with stable system prompts or repeated context, this is where the efficiency gains compound.
  • Does not cross CB-2 or Autonomy-2 safety thresholds, per Anthropic's model card. For teams with risk policies around those thresholds, Haiku 5.5 is a straightforward deployment.

What deserves a careful look:

  • The tokenizer inefficiency is real and is not disclosed in the announcement materials. You need to measure your actual token counts on representative prompts before projecting cost savings.
  • The 5× price multiplier above 100k tokens is a cliff, not a gradient. If you're building anything that occasionally handles long documents, legal text, code files, or extended conversation histories, run a histogram of your actual request lengths. The requests near the threshold determine whether you're in the "great deal" segment or the "not much better than before" segment.
  • Haiku 5.5's benchmark scores are strong for a Haiku but still below Sonnet 5.5. If you're moving workloads from Sonnet-class models to Haiku to capture the pricing, validate on quality first. The pricing cut is not free capability.
  • Long-context benchmark numbers aren't published. Anthropic didn't release performance data for tasks that hit the over-100k tier, which is exactly the regime where the pricing cliff lives. That's a meaningful omission.
Haiku 4.5 vs Haiku 5.5 vs Sonnet 5.5
Haiku 4.5Haiku 5.5Sonnet 5.5
Input price (under 100k tokens)$1.00/M$0.10/M$2.00/M
Output price (under 100k tokens)$5.00/M$0.50/M$10.00/M
Input price (over 100k tokens)N/A$0.50/MN/A (flat)
Cache reads (under 100k)—$0.01/M$0.10/M
Terminal-Bench 4.00.0%39.2%70.6%
OSWorld 2.1 offline15.7%72.4%—
HLE (no tools)10.2%45.9%—
Adjustable effortNoYesYes
GDPval-AA v2.1 (Elo)7351,6201,844
Tokenizer efficiency vs predecessorbaseline~0.8× (1.25× more tokens)—

What builders need to know

Specific steps before cutting over:

  1. Measure your tokenizer delta on representative prompts. Take 20–30 real requests from your production traffic, run them through both tokenizers, and compute the ratio. 1.25× is Simon Willison's finding on his test set. Your prompts may tokenize differently.
  2. Run a histogram of your request lengths. Plot your actual request token counts (on the Haiku 5.5 tokenizer) and find what percentage fall above and below 100k. If more than 5–10% of your volume is near-threshold, the cliff matters.
  3. Model cost in two tiers, not one. Build a spreadsheet with two rows: below-threshold cost and above-threshold cost, at realistic token counts with the new tokenizer. Then weight by your actual request-length distribution.
  4. Validate quality on your evals before cutting production. Haiku 5.5 is a significantly different model. The benchmark improvements are real, but instruction-following behavior changes with each model generation. Run your existing eval suite before flipping the switch.
  5. Check whether Sonnet 5.5 cache savings affect your model tier decision. The 50% cache-read price cut on Sonnet 5.5 ($0.20 → $0.10/M) may narrow or close the cost gap between Sonnet 5.5 and Haiku 5.5 for prompts where you can cache most of the context. Do that math too.
  6. Model ID: claude-haiku-5-5. Available on the Anthropic Platform, AWS Bedrock, Google Cloud Vertex AI, and Azure.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.