
By submitting, you consent to our use of your data. Privacy Policy.
Category
وكلاء الذكاء الاصطناعي
Share the article
Picking the strongest model available is the most expensive habit in agent engineering, and it rarely looks like a decision. It looks like a sensible default, chosen once during a prototype and never revisited, quietly applied to every step in a workflow including the ones that classify an email or reformat a date.
The bill for that habit is now public. Uber exhausted its entire 2026 AI tooling budget by April, four months into the year, across roughly 5,000 engineers, after internal Claude Code adoption climbed from 32% to 84%. Average spend landed between $150 and $250 per engineer per month, with power users between $500 and $2,000. Uber has since capped usage at $1,500 a month per person.
Microsoft reached the opposite conclusion from similar numbers. It is cancelling most internal Claude Code licences across its Experiences and Devices division, the group behind Windows, Microsoft 365, Outlook, Teams and Surface, and moving those engineers to GitHub Copilot CLI. Adoption there had reached 84% to 95% of the cohort by April. Flat seat licensing had hidden the token consumption until usage-based billing made it visible all at once.
Neither company had a model quality problem. They had a routing problem. Anthropic's current lineup spans a 10x range between its cheapest and most expensive tier, and the gap between a well-routed agent and a badly routed one is roughly that same multiple.
The current Claude lineup and what each tier is for
These are the standard API rates from Anthropic's pricing documentation, per million tokens.
Model | Input | Output | Cache read | Batch (in/out) | Built for |
|---|---|---|---|---|---|
Fable 5.1 | $10.00 | $50.00 | $0.25 | $5.00 / $25.00 | Hardest software engineering and long-horizon reasoning |
Opus 5.5 | $4.00 | $20.00 | $0.20 | $2.00 / $10.00 | Demanding reasoning, review, long-running agentic work |
Opus 5 | $5.00 | $25.00 | $0.50 | $2.50 / $12.50 | Previous flagship, now more expensive than Opus 5.5 |
Sonnet 5.5 | $2.00 | $10.00 | $0.20 | $1.00 / $5.00 | The broad production tier |
Sonnet 5 | $2.00 | $10.00 | $0.20 | $1.00 / $5.00 | Everyday building |
Haiku 4.5 | $1.00 | $5.00 | $0.10 | $0.50 / $2.50 | Speed and volume |

Two things in that table catch people out.
Opus 5 now costs more than Opus 5.5 on every line. If a config file still pins Opus 5 out of habit, it is paying a 25% premium for the older model. We covered why Opus 5.5 changed the economics when it launched.
And Sonnet 5's $2 and $10 rates were introduced as promotional pricing due to expire on 31 August 2026. The scheduled increase to $3 and $15 was cancelled, and those rates are now standard. Plans built around an expected price rise can be unwound.
One model sits outside the table. Mythos 5.1 matches Fable 5.1's $10 and $50 but is invite-only through Project Glasswing, so it is not a tier most teams can route to.
The 30% tokenizer change that hides part of the bill
Comparing sticker prices across these models understates the gap, and the reason is buried in a footnote.
Anthropic's documentation states that Claude 4.7 and later models use a newer tokenizer that "produces approximately 30% more tokens for the same text," while Sonnet 4.6 and earlier use the previous one. The company notes the exact increase depends on content and workload shape, so treat 30% as a direction rather than a constant.
Follow that through the lineup and Haiku 4.5 is the only current tier still on the older tokenizer. The same document that costs 100 tokens on Haiku can cost around 130 on Opus 5.5 or Fable 5.1 before either model reasons about it. The headline spread between Haiku and Fable is 10x on price. Measured in work completed rather than tokens billed, the practical gap is wider.
This is the same distinction we drew about cheaper models not producing cheaper agents. A rate is what you are charged per token. Your bill is that rate multiplied by tokens spent, and the second number moves independently of the first.
Which Claude model belongs on which agent step
Anthropic's own guidance is to choose "Haiku for simple tasks, Sonnet for most production workloads, and Opus for the most complex reasoning." That is correct and too coarse to route with. Here is the version that survives contact with a real workflow.
Agent step | Model | Why |
|---|---|---|
Intent classification, routing, tagging | Haiku 4.5 | Bounded output, verifiable, runs thousands of times a day |
Extraction from documents and forms | Haiku 4.5 | Schema-constrained, so errors surface immediately |
Summarisation for a human to read | Haiku 4.5 or Sonnet 5 | Move up only when tone or nuance matters |
Drafting customer-facing replies | Sonnet 5 or 5.5 | Quality is visible externally, cost still matters |
Multi-step tool use and orchestration | Sonnet 5.5 or Opus 5.5 | Needs reliable tool selection across many turns |
Code review and debugging | Opus 5.5 | Failure is expensive and hard to detect downstream |
Long-horizon autonomous work | Opus 5.5 or Fable 5.1 | Errors compound across hours of execution |
Final approval before an irreversible action | Opus 5.5 or a human | The step where being wrong actually costs something |
The rule underneath the table: spend on the steps where a mistake is expensive or hard to detect, and economise on the steps where it is cheap and obvious. A misclassified support ticket surfaces in seconds and costs almost nothing. A bad decision inside a four-hour autonomous run surfaces at the end, after it has been built on.
One caveat changes how the table should be read: the effort setting can flip which tier is cheaper. The routing above assumes the smaller model costs less per task, which holds at matched or lower effort. At maximum effort it can reverse. On Artificial Analysis's Intelligence Index, Sonnet 5.5 at max effort uses around 193,000 output tokens per task and costs $7.62 per task, while Opus 5.5 at max uses around 120,000 and costs $5.98, with a higher score. Sonnet 5.5, which became Anthropic's default model on 28 September and runs more than 30% faster than Sonnet 5 at the same $2 and $10, is excellent value at sensible effort settings. It is not the cheap option when every call runs at maximum. Set effort per step, and measure cost per task at the effort you deploy.
That principle generalises past Anthropic, which is why we wrote about assigning models per agent step as a design pattern rather than a vendor question.
The overhead you pay before any work happens
Agent workloads carry costs that never appear in a model comparison, because they are charged before the model does anything useful.
Every request that includes tools carries a tool-use system prompt. On Opus 5.5 that is 286 tokens. On Haiku 4.5 it is 496, rising to 588 when tool choice is forced. The cheap model has the higher fixed overhead, which narrows Haiku's advantage on very short calls and disappears entirely on long ones.
The toolsets are heavier again. Declaring the computer use toolset adds roughly 4,500 input tokens per request. The browser use toolset adds about 6,600, and enabling all four optional members adds another 880. The bash tool adds 325 on recent Opus models, and the text editor tool 700.
An agent doing browser work on Sonnet 5 therefore starts every single request around 7,000 tokens in the hole before reading a page. At high call volume that overhead, not the model tier, is often the largest line on the invoice. It is also invisible unless you are tracing the runs rather than reading a monthly total.
Three mechanisms claw most of it back:
Prompt caching. Cache reads cost 10% of base input, and better on the expensive tiers: 5% on Opus 5.5, and 2.5% on Fable 5.1, which is $0.25 against a $10 base. A five-minute cache pays for itself after a single read.
The Batch API. A flat 50% off input and output for anything that does not need to be synchronous. Most back-office automation qualifies and almost nobody uses it.
Effort discipline. The 1M token context window is available at standard pricing on Claude 4.6 and later, which makes it tempting to stuff the window. You are billed for what you put in it. The reasoning effort setting matters just as much: maximum effort multiplies output tokens, which is why a smaller model at max can cost more per task than a larger one at a sensible setting.
What teams actually report
The pattern practitioners keep converging on is not a single model. It is a hierarchy: a stronger model decomposes the task and dispatches subtasks to many cheap workers running in parallel, then recomposes the results. Anthropic recommends this shape directly, with Haiku as the worker tier.
The economics are the appeal. In Augment's agentic coding evaluation, Haiku 4.5 reached roughly 90% of Sonnet 4.5's performance, which is a generation-old comparison now, but it is the reason the planner-and-workers pattern keeps outperforming a single mid-tier model on both cost and wall-clock time.
The counter-signal worth taking seriously is the Uber and Microsoft experience, because both organisations had sophisticated engineers and still lost control of spend. In both cases the mechanism was the same: a default model applied uniformly, with consumption invisible until the billing model changed. Microsoft's licences were flat-rate, so nothing surfaced until they were not.
Anthropic's own worked example shows the size of the prize at the bottom of the range. Ten thousand support conversations, averaging about 3,700 tokens each, cost roughly $37 on Haiku 4.5. The same volume on Fable 5.1 would be an order of magnitude more, for work that does not need it.
How Beam handles this
In Beam, the model is a property of the node, not the workflow. Each step in a process carries its own model, with sensible defaults already set, plus automatic fallback and retry if a provider fails or a call errors.
That means a Haiku extraction step and an Opus 5.5 review step live inside the same process without anyone maintaining routing logic, and when Anthropic ships the next tier you change the steps that should change instead of re-testing the whole workflow. It is deliberately not a runtime router that guesses the best model per request. It is an explicit assignment you can read, audit and cost.
Common questions about Claude models for AI agents
Which Claude model is best for AI agents?
There is no single answer, because an agent is many steps with different failure costs. Use Haiku 4.5 for classification, extraction and other high-volume bounded work, Sonnet 5 or 5.5 for most production reasoning and tool use, and Opus 5.5 for code review, long-running autonomous work and final approvals. Fable 5.1 is reserved for the hardest engineering and long-horizon reasoning. Our Claude model overview covers the family in more depth.
How much do Claude models cost?
Per million tokens, standard API rates are: Fable 5.1 at $10 input and $50 output, Opus 5.5 at $4 and $20, Opus 5 at $5 and $25, Sonnet 5 and 5.5 at $2 and $10, and Haiku 4.5 at $1 and $5. Cache reads cost 10% of base input, or less on Opus 5.5 and Fable 5.1, and the Batch API takes 50% off both directions.
Is Opus 5.5 cheaper than Opus 5?
Yes. Opus 5.5 costs $4 input and $20 output against Opus 5 at $5 and $25, and its cache reads are $0.20 against $0.50. Opus 5.5 is both newer and cheaper, so any configuration still pinned to Opus 5 is paying more for less.
What is the difference between Sonnet 5 and Haiku 4.5 for agents?
Sonnet 5 costs twice as much per token and handles multi-step reasoning and tool use more reliably. Haiku 4.5 is built for volume and speed on well-defined tasks. Haiku also carries a higher fixed tool-use overhead, 496 tokens against Sonnet 5's 354, so its cost advantage grows with the length of the call.
Why is my Claude agent more expensive than the pricing page suggests?
Usually three reasons. Tool definitions are billed as input on every request, and the computer and browser toolsets add roughly 4,500 and 6,600 tokens respectively. Models from Claude 4.7 onward use a tokenizer producing about 30% more tokens for the same text. And reasoning output bills at the output rate, so effort settings move the bill more than the rate card does.





