LIVE
All stories ›
AI IN LIFENEWS
Tools & AppsBusiness & DealsAI ModelsResearchChips & ComputeSocietySafety & SecurityRegulation & PolicyRoboticsReviews OpenAIAnthropicGoogle & DeepMindMetaAlibaba / QwenxAIByteDance
Home › Anthropic › MODELS
MODELS

Claude Haiku 5.5 pricing: 90% cheaper under 100k tokens

Anthropic cut entry pricing to $0.10 per million input tokens, but only below 100,000 tokens — above that the rate jumps fivefold.

Claude Haiku 5.5 pricing: 90% cheaper under 100k tokens
Symbolic image: a hand turns the five-position effort dial on a compact inference rack as the status lights behind it flare brighter.

In short

Since October 7, 2026, Claude Haiku 5.5 costs $0.10 per million input and $0.50 per million output tokens for requests under 100,000 tokens — roughly 90 percent below Haiku 4.5 — but five times that above the threshold.

At a glance

  • Up to 100,000 tokens: $0.10 / $0.50 per million input/output tokens; above it, $0.50 / $2.50.
  • First Haiku-class model with adjustable effort: Low, Med, High, Xhigh, Max. Model ID: claude-haiku-5-5.
  • Benchmarks per Anthropic: OSWorld 2.1 72.4 percent, Terminal-Bench 4.0 39.2 percent, FrontierCode 1.1 46.4 percent.
  • On AWS, Google Cloud, Microsoft Azure and the Claude platform; in GitHub Copilot from October 7, 2026.
  • Simon Willison measures about 25 percent more tokens per prompt than Haiku 4.5 — a hidden surcharge.

Since October 7, 2026, Claude Haiku 5.5 costs $0.10 per million input and $0.50 per million output tokens for requests under 100,000 tokens — roughly 90 percent below Haiku 4.5 — but five times that above the threshold.

One model, two price tiers

The 100,000-token mark is the whole story. Stay under it and you pay $0.10 for input, $0.50 for output, $0.01 for cache reads and $0.125 for cache writes, each per million tokens.

Cross it and every line item multiplies by five: $0.50 input, $2.50 output, $0.05 cache reads, $0.625 cache writes. Haiku 4.5 charged a flat $1 and $5 per million tokens, according to Simon Willison.

Anthropic puts the saving against Haiku 4.5 at around 75 percent to run, and at 90 percent for requests below 100,000 tokens. Long-context workloads never see that discount.

The tokenizer nobody priced in

Willison flags a cost that appears on no rate card: Haiku 5.5 ships a less generous tokenizer and needs roughly 25 percent more tokens than Haiku 4.5 for the same prompt. The list price falls further than the invoice does.

His SVG drawing test makes the effort range concrete. At low effort the run took 7 seconds and $0.000936; at max effort the same job took 5 minutes 9 seconds and $0.033826. Haiku 4.5, a year earlier, cost $0.007583.

He also reads the tiering against GPT-6 Luna, which matches Haiku 5.5 at the standard tier and only steps up at 272,000 tokens, to $0.20 and $0.75. Above 100,000 tokens, his conclusion is that Haiku 5.5 stops being the cheap option.

Effort settings reach the Haiku tier

This is the first Haiku-class model with adjustable effort, across five steps: Low, Med, High, Xhigh and Max. The caller, not the vendor, decides where latency ends and reasoning depth begins.

The model ID is claude-haiku-5-5. It runs on the Claude platform and on Amazon Web Services, Google Cloud and Microsoft Azure.

What the scores say

Anthropic reports 72.4 percent on OSWorld 2.1 (offline subset), 39.2 percent on Terminal-Bench 4.0 and 46.4 percent on FrontierCode 1.1, plus 45.9 percent on Humanity's Last Exam without tools and 57.4 percent with them.

Elo figures round it out: 1620 on GDPval-AA v2.1 and 1578 on AA-Briefcase v1.1, with 46.4 percent on Chartography without tools. These are vendor numbers, published without independent replication at launch.

Anthropic calls Haiku 5.5 its fastest model to date and cites one customer seeing over a 30 percent latency reduction and up to 2.5x faster inference per agent turn. That is a customer report, not a measurement we verified.

Shipping in GitHub Copilot on day one

GitHub switched the model on for Copilot Pro, Pro+, Max, Business and Enterprise the same day, selectable from Visual Studio Code and the JetBrains IDEs through Xcode, Eclipse, github.com, the Copilot CLI and GitHub Mobile. Rollout is gradual, so not every account sees it yet.

Billing follows provider list pricing under usage-based billing. Business and Enterprise administrators gate access through Copilot's model policy, where new models arrive enabled unless someone turns them off.

GitHub's own early testing put Haiku 5.5 level with Claude Sonnet 5 on many coding tasks while spending, in its words, “significantly fewer tokens and steps”. The changelog attaches no figure to that.

Still unverified

None of the three pages we read states a context-window size for Haiku 5.5, and none says anything about its training data. Treat both as unknown rather than unchanged.

The open question is how the 100,000-token line behaves in real agent loops that accumulate context step by step — exactly the regime where Willison's math turns against the model.

◈ AI-GENERATED REPORT · SOURCES LINKED

FAQ

How much does Claude Haiku 5.5 cost?

$0.10 per million input and $0.50 per million output tokens below 100,000 tokens; $0.50 and $2.50 above that line.

Is Claude Haiku 5.5 cheaper than Haiku 4.5?

On list price yes — Anthropic claims about 90 percent less below 100,000 tokens. But Simon Willison measured roughly 25 percent more tokens per prompt, which eats into the saving.

Where can I use Claude Haiku 5.5?

On the Claude platform, AWS, Google Cloud and Microsoft Azure as claude-haiku-5-5, and in GitHub Copilot from the Pro plan since October 7, 2026.

Sources

More reports