8 دقيقة قراءة

Same Price, 10x the Bill: GPT-6.1 Sol vs Sonnet 5.5 vs Gemini 4 Argon

By submitting, you consent to our use of your data. Privacy Policy.

Category

وكلاء الذكاء الاصطناعي

Share the article

Nobody budgets a road trip on the price of a litre of fuel. What decides the cost of the journey is how far the car goes on each litre, and two cars filling up at the same pump can finish the trip with very different bills.

AI model pricing is now in exactly that position. Three models released or updated in the last fortnight share a list price of $2 per million input tokens and $10 per million output: OpenAI's GPT-6.1 Sol, Anthropic's Claude Sonnet 5.5 and, at its launch discount, Google's Gemini 4 Argon. On Artificial Analysis's Intelligence Index, the cheapest of the three to run costs $0.72 per task. The most expensive costs $7.62.

Same pump, same price per litre, a tenfold difference in the bill. For anyone running AI agents, which spend most of their budget on long chains of output, this is the number that matters, and it is almost never the one on the pricing page.

Cost per task on Artificial Analysis's Intelligence Index: GPT-6.1 Sol $0.72, Gemini 4 Argon $1.99, GPT-6 Astra $3.26, Claude Opus 5.5 $5.98, Claude Sonnet 5.5 $7.62

What each model actually costs per task

Artificial Analysis runs every model through the same set of tasks and records what each one spends getting there. Here are the five that most teams are choosing between this month.

Model

List price (in / out)

Cost per task

Intelligence Index

GPT-6.1 Sol

$2 / $10

$0.72

51.8

Gemini 4 Argon, launch price

$2 / $10

$1.99

52.6

GPT-6 Astra, max effort

$10 / $50

$3.26

52.7

Claude Opus 5.5, max effort

$4 / $20

$5.98

57.6

Claude Sonnet 5.5, max effort

$2 / $10

$7.62

56

The spread comes almost entirely from how many tokens each model uses to finish. At maximum effort, Sonnet 5.5 produced around 193,000 output tokens per task, the most Artificial Analysis has measured and roughly seven times what GPT-6 Astra used. Sol, at the same price per token, simply writes far less on the way to an answer.

Two details in that table are worth reading twice. GPT-6 Astra costs five times as much per token as Sol and only four and a half times as much per task, because Astra is unusually concise. And Claude Opus 5.5 at maximum effort scores higher than Sonnet 5.5 and costs less per task, which inverts the usual assumption that the smaller model is the cheaper one.

So are the cheaper models actually cheaper?

Per token, yes. Per task, it depends entirely on the model and the effort setting, and the answer can flip.

That is also why vendor claims need reading carefully. Anthropic says Sonnet 5.5 costs up to 30% less per task than Sonnet 5, and that is true: it is far more efficient than its predecessor, which used around 120,000 tokens per task to score 38. It is also, at maximum effort, the most expensive model per task in this comparison. Both statements hold, because a "cheaper per task" claim is almost always measured against the vendor's own last model, not against the market.

The effort setting moves the bill more than the model choice does in many cases. Sonnet 5.5 at lower effort costs far less than $7.62, though Artificial Analysis notes it still sits behind GPT-6 Sol on efficiency at low, medium and high. We found the same pattern in Anthropic's own data for Opus 5.5, where medium effort beat max on value. Maximum effort is a setting to reach for on purpose, not a default.

What a benchmark point is actually worth

Here is the question underneath all of this. Compared with GPT-6.1 Sol, every other model in the table buys you some extra intelligence at some extra cost:

  • GPT-6 Astra: 0.9 points more, at 4.5 times the cost per task

  • Gemini 4 Argon: 0.8 points more, at 2.8 times the cost now and 5.5 times once the launch price ends

  • Claude Sonnet 5.5 at max: 4.2 points more, at 10.6 times the cost

  • Claude Opus 5.5 at max: 5.8 points more, at 8.3 times the cost

Paying four and a half times as much for less than a point is hard to defend on almost any step. Paying eight times as much for nearly six points can be an excellent trade, on the right step.

That last qualifier is the whole answer. A composite index point is an average across many kinds of task, and your workflow is not an average. Points are worth buying on steps where a mistake is expensive or hard to detect, such as code review, a final approval, or a long autonomous run where an early error compounds. They are not worth buying on high-volume, well-bounded steps like classification or extraction, where a cheaper model's occasional error surfaces immediately and costs almost nothing.

There is also a cost that never shows up in a benchmark: proving the new model is better for you. Every switch means re-running your evaluations, checking for regressions and adjusting prompts, which is real engineering time. And public scores do not always survive contact with real work. Bloomberg reported that some Google employees with access to Gemini 4 Argon say it performs well on benchmarks but less well on some real coding tasks, a characterisation Google disputes. A two-point gain on an index is a reason to run your own evals, not a reason to migrate.

For a low-volume step, the cost of validating a switch can easily exceed a year of token savings. For a step running a hundred thousand times a month, a tenfold difference in cost per task is worth almost any amount of testing.

List prices are not a fixed reference point either

Even the sticker price is a moving target. Gemini 4 Argon's $2 and $10 is a 50% launch discount that doubles to $4 and $20, at which point its cost per task rises to $3.98, above GPT-6 Astra's. Anthropic made Sonnet 5's introductory price permanent instead of raising it on schedule. DeepSeek went the other way, raising usage fees by between 2.3 and 4.5 times, after which its revenue run-rate doubled to $1 billion without a fall in customers.

Prices move in both directions, on vendors' timelines rather than yours. Any cost model built on this month's rate card has a shelf life, which is the same lesson we drew when models got ten times cheaper and agents did not.

How to choose without guessing

Five habits make the pricing page much less misleading:

1. Measure cost per completed task on your own workload, at the effort level you actually deploy, not the one in the launch chart.

2. Set effort per step. Maximum effort on every call is the most expensive configuration available and rarely the best value.

3. Decide what a point is worth per step. Buy intelligence where errors are costly, buy efficiency where volume is high and errors are cheap to catch.

4. Budget for validation. Switch models when expected savings multiplied by volume clearly exceed the cost of re-testing, and not before.

5. Re-cost when promotions end. Put the expiry date of every launch price in the same place as the price itself.

This is the case for assigning models per agent step rather than standardising on one. In Beam, each step in a workflow carries its own model, with sensible defaults and automatic fallback and retry, so a high-volume step can move to a more efficient model without rebuilding the process, and the steps that genuinely need the extra points can keep them.

Common questions about LLM cost per task

What is cost per task, and why does it differ from price per token?

Price per token is the vendor's rate for each unit of input or output. Cost per task is what it actually costs to finish a piece of work, which depends on how many tokens the model uses along the way. Two models with the same price per token can differ tenfold per task because one reasons at far greater length than the other.

Which $2/$10 model is cheapest per task?

On Artificial Analysis's Intelligence Index, GPT-6.1 Sol costs $0.72 per task, against $1.99 for Gemini 4 Argon at its launch price and $7.62 for Claude Sonnet 5.5 at maximum effort. Sol scores slightly lower than both, at 51.8 against 52.6 and 56, so the cheapest model is not automatically the best choice for every step.

Is Claude Sonnet 5.5 expensive to run?

At maximum effort it is the most expensive model per task in this comparison, at $7.62, because it uses around 193,000 output tokens per task. At lower effort it costs much less. Anthropic reports it costs up to 30% less per task than Sonnet 5, which is accurate relative to its predecessor. At maximum effort, Claude Opus 5.5 scores higher and costs less per task.

Is a higher benchmark score worth paying more for?

It depends on the step. Extra points are worth paying for where mistakes are expensive or hard to detect, such as code review or long autonomous runs. On high-volume, well-defined steps, a cheaper and slightly less capable model usually gives better value. Validate any gain on your own workload, because public benchmark scores do not always translate to specific real-world tasks.

Will Gemini 4 Argon get more expensive?

Yes. Its $2 and $10 pricing is a 50% introductory discount. At the standard $4 and $20, Artificial Analysis puts its cost per task at $3.98, higher than GPT-6 Astra's $3.26.

ابدأ اليوم

ابدأ في بناء وكلاء الذكاء الاصطناعي لأتمتة العمليات

انضم إلى منصتنا وابدأ في بناء وكلاء الذكاء الاصطناعي لمختلف أنواع الأتمتة.

ابدأ اليوم

ابدأ في بناء وكلاء الذكاء الاصطناعي لأتمتة العمليات

انضم إلى منصتنا وابدأ في بناء وكلاء الذكاء الاصطناعي لمختلف أنواع الأتمتة.