Back

Claude Haiku 5.5: Anthropic's Cheapest, Fastest Small Model Costs About 75% Less. Prices, Benchmarks and When to Use It

On October 7, 2026, Anthropic released Claude Haiku 5.5, a small model priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens. Anthropic says it costs about 75% less to run than Haiku 4.5 on average. Here are the prices, the benchmarks, the new effort setting and where it fits.

Claude Haiku 5.5: Anthropic's Cheapest, Fastest Small Model Costs About 75% Less. Prices, Benchmarks and When to Use It
Written by
BSH Technologies
Published on2026-10-08

What Anthropic shipped

On October 7, 2026, Anthropic introduced Claude Haiku 5.5, which it calls "the cheapest, fastest, and most capable small model we've ever released." It is the third model in the Claude 5.5 generation, after Opus 5.5 and Sonnet 5.5, and it is built for high-volume, cost-sensitive work.

  • Quick, repetitive workloads: summaries, compactions, database queries and classification requests.
  • Subagent work: Anthropic says it pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding tasks.
  • Speed-sensitive tasks: it is Anthropic's fastest model to date at standard speed, so it suits live customer support and browser use.

This page is the long-form briefing behind the BSH Technologies Instagram carousel.

The price cut, in numbers

Anthropic's own pricing table, per million tokens:

  • Input: $0.10 for prompts up to 100K tokens and $0.50 above that, versus $1.00 on Haiku 4.5 and $2.00 on Sonnet 5.5.
  • Output: $0.50 up to 100K tokens and $2.50 above, versus $5.00 on Haiku 4.5 and $10.00 on Sonnet 5.5.
  • Cache reads: $0.01 up to 100K tokens and $0.05 above, versus $0.10 on Haiku 4.5.
  • Cache writes: $0.125 up to 100K tokens and $0.625 above, versus $1.25 on Haiku 4.5.

That is a 90% cut for prompts up to 100K tokens and a 50% cut above that. Anthropic says prompts up to 100K tokens made up around 90% of requests to Haiku 4.5. Its headline figure, about 75% cheaper on average, also accounts for a new tokenizer (similar to Sonnet 5.5's and Opus 5.5's) that uses slightly more tokens per task. The Claude Platform model page lists a 1M-token context window, 128K max output and a 50% Batch API discount.

As SiliconANGLE points out, the 10-cent and 50-cent rates match what OpenAI charges for GPT-6 Luna, its low-cost model launched last month.

Benchmarks: Haiku 5.5 vs. Haiku 4.5

From Anthropic's published results:

  • Computer use, OSWorld 2.1 (offline subset): 72.4% vs. 15.7% for Haiku 4.5 and 48.9% for GPT-6 Luna.
  • Agentic coding, Terminal-Bench 4.0: 39.2% vs. 0.0% for Haiku 4.5 and 16.4% for GPT-6 Luna.
  • Humanity's Last Exam: 45.9% without tools and 57.4% with tools, vs. 10.2% and 18.7% for Haiku 4.5.
  • Knowledge work, GDPval-AA v2.1: 1620 vs. 735 for Haiku 4.5 and 1437 for GPT-6 Luna.
  • Visual reasoning, Chartography: 46.4% vs. 6.4% for Haiku 4.5 and 29.1% for GPT-6 Luna.

SiliconANGLE notes Haiku 5.5 is ahead of GPT-6 Luna on all six tests where both have a score. For reference, Sonnet 5.5 still scores higher across the board, including 70.6% on Terminal-Bench 4.0.

A dial for cost vs. intelligence

Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, so developers can choose whether to optimize each task for cost or for intelligence, as they already can on Anthropic's larger models. Anthropic is clear about the limits: Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding, while Haiku 5.5 is best for narrowly scoped tasks that used to be too expensive to run at scale, like compaction, summarization and subagent work.

What customers reported

  • Asana: over 30% lower latency on task completions and up to 2.5x faster inference per agent turn in its AI Teammates evals.
  • HubSpot: 92.8% averaged over three runs on its CRM eval suite, the best score it has seen from a smaller model.
  • AlphaSense: 0.84 vs. 0.76 for Haiku 4.5 across 400 Ask in Document queries.
  • Box: 11 points higher than Haiku 4.5 at about half the latency in early testing.

Also announced on October 7

  • Sonnet 5.5 cache reads cut 50%, from $0.20 to $0.10 per million tokens, which Anthropic says makes Sonnet 5.5 about 20% cheaper on most agentic work.
  • Monthly API credits for Claude Platform: $100 for Max 5x, $200 for Max 20x and up to $500 pooled for Team subscribers, usable on any Claude model.
  • SDK updates: the Claude Python and TypeScript SDKs add computer use and browser use support in beta.
  • Availability: Haiku 5.5 is live on the Claude Platform as claude-haiku-5-5 and on Amazon Web Services, Google Cloud and Microsoft Azure. The AWS announcement covers Amazon Bedrock availability with Regional data residency.

Safety notes

Anthropic says Haiku 5.5 shows far fewer instances of misaligned behavior than Haiku 4.5. Its cybersecurity safeguards permit a wider range of defensive tasks than Sonnet 5.5's but still block penetration testing and other attacker-leaning techniques; its biology safeguards match Sonnet 5, Sonnet 5.5 and Opus 5.

What this means for teams

  • High-volume pipelines: classification, tagging, summarization and support triage that were borderline on cost are now worth re-pricing at $0.10 / $0.50.
  • Agent builders: use a larger model as the planner and Haiku 5.5 as the fast subagent for lookups, extraction and compaction.
  • Budget owners: keep prompts under 100K tokens where you can, since that is where the 90% cut applies, and test the effort setting before defaulting to a bigger model.

Primary sources

  • Anthropic: Introducing Claude Haiku 5.5 (Oct 7, 2026)
  • Claude Platform docs: Claude Haiku 5.5 model overview and pricing
  • AWS Machine Learning Blog: Introducing Claude Haiku 5.5 on AWS (Oct 7, 2026)
  • SiliconANGLE: Anthropic releases Claude Haiku 5.5 small model and halves Sonnet 5.5 cache read prices (Oct 7, 2026)

How BSH can help

At BSH Technologies we help teams pick the right model for each job and wire it into production: routing simple tasks to fast, cheap models, keeping complex work on larger ones and measuring cost and quality as you go. If you want to cut your AI bill without cutting results, our Thrissur engineers can help you design the routing and evaluation layer.

Frequently asked questions

How much does Claude Haiku 5.5 cost?

For prompts up to 100,000 tokens, $0.10 per million input tokens and $0.50 per million output tokens. Above 100,000 tokens, $0.50 and $2.50. Haiku 4.5 cost $1 and $5.

Is Claude Haiku 5.5 really 75% cheaper?

Anthropic says it costs around 75% less to run on average. List prices are 90% lower up to 100K tokens and 50% lower above that, and the average accounts for a new tokenizer that uses slightly more tokens per task.

When should I use Haiku 5.5 instead of Sonnet 5.5?

For narrowly scoped, high-volume or speed-sensitive tasks such as summaries, classification, compaction, customer support and subagent work. Anthropic still recommends Sonnet 5.5 or Opus 5.5 for complex agentic coding.

Related Topics

#Anthropic#Claude#Claude Haiku 5.5#AI Pricing#LLM#AI Agents#Small Models

From the blog

View all posts
OpenAI Released 722 AI-Written Math Papers on GitHub: What's Inside, What's Verified and Why Mathematicians Are Split
blog.categories.ai

OpenAI Released 722 AI-Written Math Papers on GitHub: What's Inside, What's Verified and Why Mathematicians Are Split

On October 6, 2026, OpenAI published 722 mathematical manuscripts in 372 result families, produced by an unreleased internal model, to a public GitHub repository with Lean formalizations for many of the proofs. Here is what was released, what has actually been checked, and why the math community is divided.

BSH Technologies
BSH Technologies · 2026-10-07
OpenAI Is Watermarking ChatGPT Text in the EU: How textGrain Works and What It Can't Prove
blog.categories.ai

OpenAI Is Watermarking ChatGPT Text in the EU: How textGrain Works and What It Can't Prove

On October 5, 2026, OpenAI said it will add an invisible text watermark called textGrain to eligible ChatGPT and Codex output in the European Union to meet the EU AI Act, let API customers worldwide opt in for select models, and open its detector to approved researchers only.

BSH Technologies
BSH Technologies · 2026-10-06