8 Min. Lesezeit

What Can Gemini 4 Argon Do With a Million Tokens?

By submitting, you consent to our use of your data. Privacy Policy.

Kategorie

KI-Agenten

Artikel teilen

In early 2024, asking a frontier model for a long answer meant getting a few pages back. Most of them stopped writing at somewhere between 4,000 and 8,000 tokens, and anything longer had to be stitched together from several requests, usually with a seam showing.

The ceiling has been rising quietly ever since, and last week it jumped. Gemini 4 Argon can now produce up to one million tokens in a single response, up from 64,000 on Google's previous generation. OpenAI's GPT-6 family tops out at 128,000. So the obvious question is a fun one: what would you actually do with a million tokens of output?

The honest answer is that most jobs will never need it. The ones that do are interesting, and a surprising number of them are agent jobs.

How much is 1 million tokens, really?

Some quick arithmetic, using the common rule of thumb that a token is about three quarters of an English word:

  • About 750,000 words. That is longer than War and Peace.

  • Roughly 1,500 single-spaced pages of plain text.

  • About $10 for a response that uses the full million at Argon's launch price of $10 per million output tokens, and about $20 at the standard price.

That last line matters more than it looks. A maxed-out response is not free, and it is also one very large thing to check.

For a sense of what Google itself has pointed Argon at: agents running on it have freed more than 300 TiB of memory across Google's data centres, migrated C and C++ code to Rust in codebases of up to 800,000 lines, and rewritten 32,000 lines of performance-critical code in a video decoder to run 2.7 times faster with identical output. One quantum computing subroutine came back 40% more efficient than the published baseline, in minutes.

Four everyday jobs that suddenly fit in one pass

These are the tasks where a 128,000-token cap forces you to chop the work up, and where chopping it up causes real problems: inconsistent terminology, lost context, or a seam in the middle of a document.

1. Translate an entire contract pack. A 300-page set of agreements is around 150,000 words, or roughly 200,000 tokens of output. That is past every other flagship's cap, but well inside Argon's. The win is not speed. It is that a defined term means the same thing on page 280 as it did on page 3.

2. Write every product description in a catalogue. Ten thousand items at 50 words each is around 500,000 words, or 670,000 tokens. In one pass, the tone, length and structure stay consistent across the lot, which is exactly what goes wrong when the job is split between twenty requests.

3. Rewrite the employee handbook in plain English. A 200-page policy handbook is roughly 133,000 tokens, just over the 128,000 line. It is a small overrun with a big practical consequence: one coherent document instead of two halves edited separately.

4. Document a whole codebase. Reference documentation for a large repository, or a full test suite, can run to hundreds of thousands of tokens. Doing it in a single pass keeps naming, cross-references and conventions consistent from the first module to the last.

Four agent jobs where the ceiling really earns its keep

This is where a long output stops being a party trick. Agents do repetitive, structured work at volume, and the output of that work, written up properly, gets long fast.

5. Screen a thousand candidates, with reasons. A screening agent that scores applicants is useful. One that writes a short, criterion-by-criterion rationale for every decision is auditable. At 300 words per candidate, a thousand applicants comes to around 400,000 tokens. This is the kind of high-volume work Beam's Candidate Screening AI Agent is built for, and for recruiting teams the HR and RPO use case is where we see it most.

6. Explain every exception at month-end. In finance, the number is rarely the hard part. The explanation is. An agent writing a short note for each of 2,000 reconciling items, or every flagged line from a month of supplier invoices, produces something like 250,000 tokens of commentary. That sits naturally alongside an invoice processing agent that has already done the matching.

7. Migrate a codebase in one coherent sweep. Google's own example is the strongest one here. Moving hundreds of thousands of lines from one language to another works better when the model can keep the whole design in view and write the result without forgetting decisions it made earlier. Even then, Google says those rewrites go through automated and manual auditing, emulation testing and review before reaching production.

8. Audit the agents themselves. Picture an oversight agent whose job is to read a day of runs from fifty other agents and write the complete exception report: what each agent did, where it escalated, which decisions need a person to look again, with every item traced to the run it came from. Two thousand runs at 150 words each is around 400,000 tokens. This is the heaviest job on the list, and the one where a long output is most useful, because the value lies in seeing everything at once.

The eight jobs at a glance

Job

Rough output

Fits a 128K cap?

Handbook rewritten in plain English

~133,000 tokens

Just over

Contract pack translated

~200,000 tokens

No

Month-end exception commentary

~250,000 tokens

No

Candidate screening with reasons

~400,000 tokens

No

Oversight report across 50 agents

~400,000 tokens

No

Product catalogue descriptions

~670,000 tokens

No

Full codebase documentation

Hundreds of thousands

No

Large codebase migration

Hundreds of thousands

No

These are estimates at about 0.75 words per token, meant to show scale rather than precise sizes. The pattern is the useful part: almost every job on the list is a volume job, and most of them are the kind of structured, repeatable work agents already do.

Estimated output for eight jobs, from 133K to 670K tokens, against the 128K GPT-6 cap and Gemini 4 Argon's 1 million token ceiling

Should an agent ever write a million tokens at once?

Mostly, no, and that is the most useful thing to know about this feature.

A single enormous response has three practical problems. It is expensive, at up to $20 a time at standard pricing. It is hard to review, because a mistake on page 900 is as costly as one on page 9 and much harder to find. And if something goes wrong late in the run, you redo the whole thing. Early users also report that Google's own coding harness compacts context at around 250,000 tokens, so the model's ceiling and the tool's ceiling are not yet the same number.

For most agent work, the better design is the opposite of one giant output: a sequence of steps, each producing something reviewable, with checks between them and an independent record of what happened. That is how Beam workflows run. Each step carries its own model, defined actions and escalation paths, and only the steps that genuinely need a long, coherent output, such as a full translation, a migration or an end-of-day audit report, should be assigned a long-output model.

So the million-token ceiling is less a new way to work and more a constraint removed. The jobs that used to need awkward splitting can now be done whole. Everything else should still be done in steps.

Common questions about the 1 million token output limit

What does a 1 million token output limit mean?

It means a model can produce up to one million tokens in a single response, rather than stopping and needing to be prompted again. Gemini 4 Argon is the first major model to offer this, up from 64,000 tokens on Google's previous generation. OpenAI's GPT-6 models cap output at 128,000 tokens.

How many words is 1 million tokens?

Roughly 750,000 words in English, using the common estimate of about 0.75 words per token. That is around 1,500 single-spaced pages, longer than War and Peace. The exact figure depends on the language and content, and code typically uses more tokens per line than prose.

How much does a 1 million token response cost on Gemini 4 Argon?

At Argon's introductory price of $10 per million output tokens, a response that uses the full million costs about $10 in output alone, plus input. At the standard price of $20 per million output tokens, it costs about $20. Most real tasks use far less: Artificial Analysis measured Argon at around 62,000 output tokens per task on average.

Can I use Gemini 4 Argon's 1 million token output now?

Only through Google's Fairwind Program, which covers vetted Google Cloud customers, government agencies and cybersecurity partners. Paid API customers and Google AI Ultra subscribers are next, but Google has not published dates.

Do AI agents need a 1 million token output?

Rarely. Most agent work is better split into steps with checks between them, which is cheaper and easier to review. A long output earns its place on specific steps where splitting would break consistency, such as translating a large document, migrating a codebase, or writing a single audit report across many agents' work.

Heute starten

Starten Sie mit KI-Agenten zur Automatisierung von Prozessen

Nutzen Sie jetzt unsere Plattform und beginnen Sie mit der Entwicklung von KI-Agenten für verschiedene Arten von Automatisierungen

Heute starten

Starten Sie mit KI-Agenten zur Automatisierung von Prozessen

Nutzen Sie jetzt unsere Plattform und beginnen Sie mit der Entwicklung von KI-Agenten für verschiedene Arten von Automatisierungen