
By submitting, you consent to our use of your data. Privacy Policy.
Kategorie
KI-Agenten
Artikel teilen
Open the trace of a well-built agent run and the shape is usually the same. One large model sits at the top, reads the request and writes a plan. Below it, a fan of smaller calls each takes one narrow job: pull a figure from a filing, classify a ticket, summarise a thread. Then the large model reads what came back and writes the answer.
Anthropic released Claude Haiku 5.5 on 7 October and priced it for the bottom of that fan. It costs $0.10 per million input tokens and $0.50 per million output for prompts up to 100,000 tokens, a tenth of Haiku 4.5's $1 and $5. Anthropic pitches it for summaries, classification and compaction, and says it "pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work." One launch customer, Rogo, describes the pattern directly: while a bigger model builds a deck, a Haiku 5.5 subagent goes into the 10-K and pulls the segment revenue line the deck needs.
Our view: at this price, the subagent call stops being the thing to optimise. What decides the bill now is the orchestrator's own tokens, the cost of every hand-off, a price step at 100,000 tokens that is easy to cross by accident, and how often the cheap step gets it wrong.
In short:
What it is: Anthropic's cheapest current model and its fastest at standard speed, with a 1M-token context window, 128K max output and adjustable effort.
Haiku 5.5 price: $0.10 input and $0.50 output per million tokens for prompts up to 100,000 tokens. Over that, every rate is five times higher: $0.50 and $2.50.
Benchmarks: large gains over Haiku 4.5 on Anthropic's numbers (vendor claim). Artificial Analysis scores it 43 at maximum effort and 34 at the default medium effort, against 17 for Haiku 4.5.
The subagent math: in our illustrative run, 20 Haiku 5.5 subagents are 7% of the bill. The Opus 5.5 orchestrator and its hand-offs are the other 93%.
What to do: keep subagent prompts under 100,000 tokens, keep briefs and answers short, and spend the savings on checking the cheap step.
Claude Haiku 5.5 price and specs at a glance
Prices are standard API rates per million tokens from Anthropic's pricing documentation, checked 8 October.
Haiku 5.5 (prompt up to 100k / over 100k) | Haiku 4.5 | Sonnet 5.5 | Opus 5.5 | |
|---|---|---|---|---|
Input | $0.10 / $0.50 | $1.00 | $2.00 | $4.00 |
Output | $0.50 / $2.50 | $5.00 | $10.00 | $20.00 |
Cache read | $0.01 / $0.05 | $0.10 | $0.10 | $0.20 |
Batch (in / out) | $0.05 / $0.25, then $0.25 / $1.25 | $0.50 / $2.50 | $1.00 / $5.00 | $2.00 / $10.00 |
Context window | 1M | 200k | 1M | 1M |
Max output | 128K | 64K | 128K | 128K |
Default effort | Medium | n/a | High | Medium |
Haiku 5.5 uses the newer tokenizer from Claude 4.7 onward, so the same text counts as roughly 30% more tokens than on Haiku 4.5. Anthropic accounts for that and still puts the average saving at around 75% per piece of work. Priority Tier is not supported.
It is available on the Claude API, Amazon Bedrock, Google Cloud's Vertex AI, Microsoft Foundry and Claude Platform on AWS. The same launch halved Sonnet 5.5's cache-read price to $0.10.
What the Haiku 5.5 benchmarks show
Anthropic's launch table shows a model that has moved a long way from its predecessor. Haiku 5.5 scores 39.2% on Terminal-Bench 4.0 against 0.0% for Haiku 4.5, 72.4% on the offline subset of OSWorld 2.1 against 15.7%, and 45.9% on Humanity's Last Exam without tools against 10.2% (vendor claim). Sonnet 5.5 still leads on all three, at 70.6%, 83.9% and 56.9%.
The independent numbers point the same way. Artificial Analysis scores Haiku 5.5 at 43 on its Intelligence Index at maximum effort, second of 182 models in its price class, at $0.21 per index task. At the default medium effort it scores 34 for $0.05 per task. Haiku 4.5 with reasoning scores 17 and costs $0.28 per task.
Effort is the dial to watch. At maximum, Haiku 5.5 generated 440 million output tokens to run the index, against 54 million at medium. OpenAI's GPT-6 Luna lists the same $0.10 and $0.50 price and scores 38 at maximum effort for $0.07 per task, so Haiku 5.5 at max buys five more points for three times the cost. On a high-volume subagent step, medium or high effort is usually the better trade.
Anthropic is clear about the ceiling: Sonnet 5.5 and Opus 5.5 "remain better choices for complex agentic coding tasks," while Haiku 5.5 suits "narrowly scoped tasks" like subagent work.
The subagent call is no longer the expensive part
Take one illustrative run, with every assumption stated. An Opus 5.5 orchestrator caches a 15,000-token system prompt, reads a 2,000-token task, writes a 3,000-token plan and 20 briefs of 250 tokens each. Twenty Haiku 5.5 subagents each read 8,000 tokens and write 1,500, including thinking. The orchestrator then reads 20 answers of 600 tokens and writes a 4,000-token synthesis.
Each subagent call costs $0.00155. Writing its brief costs the Opus 5.5 orchestrator $0.005 in output tokens, and reading its answer costs another $0.0024 in input. The hand-off around the job costs almost five times as much as the job.

With an Opus 5.5 orchestrator, passing work to a Haiku 5.5 subagent and back costs 4.8 times the work itself.
Add it up and the whole run costs about $0.445. The 20 subagents account for $0.031 of that, or 7%. The orchestrator's planning and synthesis cost $0.266, and the hand-offs $0.148. Put every step on Opus 5.5 instead and the subagents alone cost $1.24, for a total of $1.65.

Moving the volume work to Haiku 5.5 cuts the run by almost three-quarters, and leaves the orchestrator as the main cost.
It is the lesson of our piece on how models at the same price produce a tenfold spread in cost per task, seen from the other side: once one line gets very cheap, the remaining lines decide the bill.
Where the orchestrator's money goes, and how to cut it
Swap the orchestrator to Sonnet 5.5 and the same run costs $0.238, because every Sonnet rate is half of Opus 5.5's, including the newly halved cache read. Whether Sonnet plans well enough for your task is a question for your evaluations.
Our guide to choosing between Opus, Sonnet and Haiku covers which tier suits which step. For this pattern, three habits matter more than the tier:
Short briefs. A brief is orchestrator output, the most expensive token in the run.
Structured answers. Ask subagents for a fixed schema with only the fields the orchestrator needs. A 300-token answer costs half as much to read as a 600-token one.
Cache the stable parts. The orchestrator's system prompt and tools repeat on every turn. Cache reads cost a twentieth of base input on both Opus 5.5 and Sonnet 5.5.
The orchestration pattern matters too. Passing results to a cheap aggregator keeps the expensive model from reading every intermediate answer.
Why the 100,000-token line matters for subagents
Haiku 5.5 is priced by prompt length. Up to 100,000 tokens, input costs $0.10 per million. Past that, input, output, cache writes and cache reads all cost five times as much, and the step applies to the whole request, not only the tokens above the line.
That matters because subagents are often the steps that receive big contexts. Hand a subagent a whole long document instead of the relevant section and a call that would have cost $0.00155 at 8,000 tokens costs about $0.064 at 120,000. Twenty of those come to $1.28, about three times the rest of our illustrative run.

A subagent that receives 120,000 tokens instead of 95,000 costs about six times as much per call.
Even over the line, Haiku 5.5 is cheaper than Sonnet 5.5, which would charge $0.255 for the same call. But the step changes how you feed subagents: retrieve the relevant passages, split long documents across calls, and keep each prompt comfortably under 100,000.
Two details need care. Anthropic's documentation does not say whether cached tokens count toward the 100,000, so assume they do until your bills say otherwise. And Haiku 5.5 keeps previous thinking blocks in context by default, so a subagent that runs a long multi-turn conversation can drift over the line without anyone sending it a large document.
The cheap step's error rate is the real constraint
When a subagent call costs a fraction of a cent, the expensive thing is a wrong answer nobody catches. A misread revenue figure flows into the synthesis, and from there into a decision.
The good news is that checking is now cheap as well. In our illustrative run, a second Haiku 5.5 call that verifies every subagent answer adds $0.031, about 7% of the run. Re-running one job in twenty on Opus 5.5 when a check fails adds $0.062. Having the Opus orchestrator re-read every source document to check its subagents would cost more than the subagents themselves.
So spend the savings on verification: schema checks on every answer, a second cheap call on the fields that matter, and escalation to a stronger model only when a check fails. Plan for refusals too, since Haiku 5.5's safety classifiers can decline a request with no server-side fallback. And trace every run, because a fan of twenty calls hides failures that a monthly total never shows.
How to set this up per step
The pattern Anthropic describes is a per-step decision: a stronger model for the plan and the final judgement, a cheap one for the volume.
That is how assigning models per agent step works in Beam. The model is a setting on each step of an agent, and each step starts from a default drawn from the workspace's preferred models. Steps can carry fallback models if a call fails, and can retry when accuracy drops below a threshold. There is no runtime router guessing which model to use. The assignment is explicit, so each step's cost can be read and audited, and trying Haiku 5.5 on one step leaves the rest of the agent untouched. Test it on your own tasks, at the effort you will deploy, before moving anything.
Common questions about Claude Haiku 5.5
When was Claude Haiku 5.5 released?
Anthropic released Claude Haiku 5.5 on 7 October 2026. It is available on the Claude API as claude-haiku-5-5, on Amazon Bedrock, Google Cloud's Vertex AI, Microsoft Foundry and Claude Platform on AWS. Anthropic commits to keeping it available until at least 7 October 2027.
What is the Haiku 5.5 price?
For prompts up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens, $0.50 per million output tokens and $0.01 per million cached input tokens. For prompts over 100,000 tokens, every rate is five times higher: $0.50 input, $2.50 output and $0.05 cached. The Batch API takes 50% off input and output.
How does Haiku 5.5 do on benchmarks?
Anthropic reports 39.2% on Terminal-Bench 4.0 and 72.4% on the OSWorld 2.1 offline subset, against 0.0% and 15.7% for Haiku 4.5 (vendor claims). Artificial Analysis scores it 43 on its Intelligence Index at maximum effort and 34 at the default medium effort, against 17 for Haiku 4.5 with reasoning.
Is Haiku 5.5 better than Haiku 4.5?
On published numbers, yes, by a wide margin, and it is much cheaper. It also has a 1M-token context window instead of 200k, 128K max output instead of 64K, and adjustable effort. It counts the same text as about 30% more tokens, uses adaptive thinking, and does not support Priority Tier, so migrations need retesting.
Should Haiku 5.5 run my agent's subagents?
It suits narrow, high-volume subagent jobs such as extraction, lookups, classification and summaries. Anthropic itself recommends Sonnet 5.5 or Opus 5.5 for complex agentic coding. Keep each subagent prompt under 100,000 tokens, ask for short structured answers, and check its outputs before the orchestrator relies on them.





