9 min leer

Gemini 4 Argon Knows When It Doesn't Know

By submitting, you consent to our use of your data. Privacy Policy.

Categoría

Agentes de IA

Compartir artículo

Less than a year ago, Google's flagship model had one of the worst habits a model can have. On Artificial Analysis's knowledge benchmark, when Gemini 3 Pro did not know the answer to a question, it guessed wrong 88% of the time instead of saying so. It knew a great deal. It simply would not admit what it didn't.

Google's new flagship has nearly inverted that. Gemini 4 Argon, announced on 30 September, has a hallucination rate of 15% on the same independent benchmark, the lowest of any model at its level of intelligence. GPT-6 Astra sits at 51%. GPT-6.1 Sol sits at 54%.

Most of the launch coverage led with the benchmark table and the one-million-token output limit. Both are real. For anyone deploying agents, the calibration number is the one that changes how a model behaves in production, and it comes with two catches that the headlines mostly skipped.

What Gemini 4 Argon is

Argon is Google DeepMind's new frontier model, built for coding, enterprise knowledge work and cyber defense, and the successor to the Gemini 3.8 generation. Its most unusual specification is output: a single response can now run to one million tokens, up from 64,000 on the previous generation. That is aimed at work that used to need handoffs between segments, such as large code migrations and long multi-step research.

Almost nobody can use it yet. Argon is live only for members of Google's Fairwind Program, a group of vetted Google Cloud customers, government agencies and cybersecurity partners. Paid API customers and Google AI Ultra subscribers are next, with no published date, followed by broader access "as soon as possible."

Pricing launches at $2 per million input tokens and $10 per million output, but that is a 50% introductory discount. The standard rate afterwards is $4 and $20, identical to Claude Opus 5.5. Cached input is 95% cheaper than standard.

The benchmarks, read honestly

On the numbers Google published, Argon leads most of the field, often clearly. It also trails in a consistent place.

Benchmark

Gemini 4 Argon

Claude Opus 5.5

GPT-6 Astra

DeepSWE v1.1 (long software work)

77.9%

74.2%

74.1%

AutomationBench

51.3%

42.5%

41.4%

GraphWalks 256K to 1M (long context)

84.2%

66.8%

71.8%

Terminal-bench 4.0

57.4%

66.4%

58.2%

FrontierSWE v2

55.0%

62.3%

65.5%

OSWorld-2.0 (computer use)

69.2%

not reported

72.6%

The pattern is worth reading rather than counting. Argon is strongest on long, structured work over a lot of context, and weakest at driving a terminal or a desktop, where Opus 5.5 and Astra both pull ahead. Its two software benchmarks also tell different stories: 77.9% on DeepSWE, where it leads, and 55.0% on FrontierSWE, where it trails Astra by more than ten points. Which one resembles your codebase matters more than either headline.

Two caveats apply to that whole table. These are Google's own results, there is no technical report yet, and no outside lab has replicated them. On the independent side, Artificial Analysis puts Argon level with GPT-6 Astra on its Intelligence Index, at around 53 each, which is a strong result rather than a clear lead. On the same index, Claude Opus 5.5 at maximum effort scores 57.6 and Claude Sonnet 5.5 scores 56, so Argon is a frontier contender rather than the new leader.

Google's internal examples are striking and should be read as Google's: agents running on Argon found memory optimisations across its data centres that freed more than 300 TiB, with a projected total between 500 TiB and 1 PiB, and handled codebases above 800,000 lines.

Why the hallucination rate matters more than the benchmarks

The number needs reading carefully, because it is easy to overstate. On Artificial Analysis's AA-Omniscience benchmark, the hallucination rate is the share of wrong answers among all the questions a model did not get right. The rest are cases where it declined. Models answer without tools or context, and are told that abstaining beats guessing.

Gemini 4 Argon vs GPT-6 Astra on AA-Omniscience: accuracy 50% vs 63%, hallucination rate 15% vs 51%

So a 15% rate means that when Argon does not know something, it says so 85% of the time. Astra, at 51%, guesses wrong about half the time it is out of its depth.

That is not the same as Argon knowing more. On the same benchmark its accuracy is 50%, against 63% for Astra. Argon knows less, and is far better at noticing the edge of what it knows.

For an agent, that trade usually runs in Argon's favour. Raw recall matters less inside an agent, because a well-built agent retrieves facts from documents, systems and tools rather than from memory. What it cannot retrieve is judgement about when it is uncertain. A model that guesses confidently produces an answer that looks exactly like a correct one, and in a multi-step workflow that wrong answer becomes the input to every step after it.

A model that stops and says it does not know produces a visible failure instead of a silent one. Visible failures are the ones you can route to a human, retry, or fix.

The 1M output limit, and the bill attached to it

The output ceiling and the second catch are the same fact viewed from two directions.

Argon thinks at length. Artificial Analysis measured it averaging 62,000 output tokens per task, against 27,000 for GPT-6 Astra. At the introductory price that still makes it cheaper per task than Astra: $1.99 against $3.26. At the standard price, Argon costs $3.98 per task, which is more than Astra. And OpenAI's GPT-6.1 Sol, released the day before at the same $2 and $10 list price, costs $0.72 per task on the same index, less than half of Argon even at the launch discount.

That is the gap between a rate and a bill that we keep returning to with cheaper models and agent costs. Argon is 60% cheaper per token than Astra at standard pricing and around 22% more expensive per task, because it spends more than twice the tokens getting there. Any cost case built on the launch price will need rebuilding when the discount ends.

The one-million-token ceiling is also only as useful as the tool around it. Early users report that Google's own coding harness compacts the context at roughly 250,000 tokens, well below what the model itself supports.

What people are saying about Gemini 4 Argon

The early reaction separates the model from the company shipping it, and is noticeably kinder to the first.

The model gets genuine respect. One widely shared account describes Argon attaching a debugger to a GPU driver and reverse-engineering a kernel interface to fix a hardware problem, unprompted. Developers broadly rate it alongside Opus 5.5 on complex work, sometimes a little behind.

The criticism lands on delivery. The most common complaint is Google's coding harness, including that compaction ceiling and the lack of a safe automatic mode. The biggest stated barrier to business use is not quality at all. It is the fear that a policy violation can suspend an entire Google account, taking email and documents with it.

There is also an internal disagreement worth knowing about. Bloomberg reported that some Google employees with direct access say Argon performs well on benchmarks but less well on real work, struggling with certain coding tasks including front-end design. Google told Bloomberg that characterisation is inaccurate, and another employee described a "large consensus" internally that the model is at the frontier. All three positions can be true at once, and the practical reading is the usual one: benchmark results are a reason to test, not a reason to switch.

Where Gemini 4 Argon fits in an agent stack

You cannot route work to Argon yet, which makes this a planning question rather than a deployment one. When access opens, the evidence points to a specific set of steps rather than a wholesale switch, the same logic behind assigning models per agent step.

Argon's profile suits long-context synthesis across large document sets, large code migrations where the output ceiling removes handoffs, and any step where a confident wrong answer is expensive and an honest "I don't know" is cheap to handle. Keep terminal and desktop control on Opus 5.5 or GPT-6 Astra, where Argon trails, and re-cost every step at the standard price rather than the launch discount.

In Beam, a model is a property of each step in a workflow, with sensible defaults plus automatic fallback and retry. When Argon becomes available, the steps that suit it can move to it without rebuilding the workflow, and the ones that do not can stay where they are.

Common questions about Gemini 4 Argon

Can I use Gemini 4 Argon now?

Only through Google's Fairwind Program, which covers vetted Google Cloud customers, government agencies and cybersecurity partners. Paid API customers and Google AI Ultra subscribers are next, and broader access follows, but Google has not published dates for either stage.

How much does Gemini 4 Argon cost?

At launch it costs $2 per million input tokens and $10 per million output, which is a 50% introductory discount. The standard price afterwards is $4 and $20, the same as Claude Opus 5.5, with cached input 95% cheaper. Because Argon uses more output tokens per task, its cost per task at standard pricing is higher than GPT-6 Astra's, and even at the launch price it costs more than twice as much per task as GPT-6.1 Sol, which has the same list price.

Is Gemini 4 Argon better than GPT-6 Astra and Claude Opus 5.5?

On Google's reported benchmarks it leads on long software work, automation and long-context tasks, and trails on terminal and desktop control. Independently, Artificial Analysis rates it level with GPT-6 Astra on overall intelligence, with Claude Opus 5.5 scoring higher. It is a strong contender rather than a clear winner, and its standout advantage is calibration rather than raw capability.

What does Gemini 4 Argon's 15% hallucination rate mean?

On Artificial Analysis's AA-Omniscience benchmark, it means that when Argon does not know an answer, it gives a wrong one 15% of the time and declines the other 85%. GPT-6 Astra gives a wrong answer 51% of the time in the same situation. Argon's overall accuracy is lower, at 50% against Astra's 63%, so it knows less but is much better at recognising when it doesn't.

What is the 1 million token output limit useful for?

It lets a single response cover work that previously had to be split across several, such as large code migrations or long multi-step research, without handing context between segments. In practice the limit is only as useful as the surrounding tooling, and early users report Google's own coding harness compacting context well below it.

Empieza hoy

Empezar a crear agentes de IA para automatizar procesos

Únase a nuestra plataforma y empiece a crear agentes de IA para diversos tipos de automatizaciones.

Empieza hoy

Empezar a crear agentes de IA para automatizar procesos

Únase a nuestra plataforma y empiece a crear agentes de IA para diversos tipos de automatizaciones.