
By submitting, you consent to our use of your data. Privacy Policy.
Category
AI Agents
Share the article
About one in twenty. That is roughly how much of the newest large model out of Europe does any work on a given token: around 50 billion of its 1.05 trillion parameters. The other 95% sit in memory, waiting to be picked for a token that needs them.
The model is Mistral Large 4, nicknamed "le Chonk", which Mistral put into public preview on its own API on 6 October. It takes text and images, Mistral's model card lists a one-million-token context window, and Mistral trained it from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters. The weights, Mistral says, drop by the end of this month.
For a CTO, the useful question is where it fits. Our view: Mistral Large 4 earns a place when you need to run a strong model under your own control, ideally in Europe. On today's independent numbers it sits behind the strongest open models, and its low price per token turns into a middling cost per task.
In short:
What it is: a mixture-of-experts model with about 1.05 trillion total and roughly 49 to 52 billion active parameters, image input, and a public preview API on Mistral Studio.
Price: $1.36 input and $4.18 output per million tokens at list, currently shown at half that ($0.68 and $2.09) during the preview.
Benchmarks: Artificial Analysis scores the preview 38 on its Intelligence Index, behind GLM-5.3 (45) and Kimi K3 (44). LMArena has no entry yet.
Open weights: promised by the end of October. The licence terms are not public yet.
What to do: test it on the specific agent steps where control matters, using your own tasks, before moving anything.
Mistral Large 4 specs and pricing at a glance
Mistral's announcement and its model card disagree slightly on two numbers, so both are shown where they differ.
Mistral Large 4 (preview) | |
|---|---|
Architecture | Granular mixture of experts |
Total parameters | 1.05 trillion (model card); "1 trillion" in the announcement |
Active parameters per token | 49 billion (announcement); 52 billion (model card) |
Vision | 1.6B-parameter vision encoder, text and image in, text out |
Context window | 1M tokens on Mistral's card; Artificial Analysis lists 524k for the preview API |
List price (per M tokens) | $1.36 input, $0.14 cached input, $4.18 output |
Preview price shown now | $0.68 input, $0.07 cached input, $2.09 output |
Training | 3,800 NVIDIA Grace Blackwell GPUs, Mistral's own datacenters in Europe |
Languages | 160+, including every official EU language |
Where to use it | Preview API on Mistral Studio |
Open weights | "By the end of the month" (Mistral); licence not yet published |
The model card shows the list price struck through with the preview price beside it, and gives no end date for the discount. Budget on the list price.
What 49 billion active out of a trillion means for speed and cost
A mixture-of-experts model splits its parameters into many specialist blocks, called experts, and a small router chooses a handful of them for each token. The model gets the knowledge capacity of a very large network while paying the compute bill of a much smaller one, because only the chosen experts run.

Speed and per-token compute follow the few experts that wake up; the memory you need to host the model follows all of them.
That design shows up in the speed numbers. Artificial Analysis measured the preview at 116 output tokens per second, which it rates as notably fast. Qwen3.8 2.4T A95B, which activates about 95 billion parameters, runs at 39. Serving hardware differs between providers, so treat this as direction rather than a law.
There are two catches. First, a mixture-of-experts model is cheap to run per token but heavy to host, since every expert has to sit in memory whether or not it is used. Self-hosting 1.05 trillion parameters is a multi-GPU commitment, which matters for anyone planning to take the open weights in-house. Second, speed per token tells you nothing about how many tokens a model needs to finish a job.
What the independent benchmarks say so far
Mistral's launch post is full of strong results. It reports 61.7% on DeepSWE v1.1, 28.3% on Terminal-Bench 4 and 59.9% on AutomationBench (vendor claim). In a blind coding evaluation it ran with Surge AI, annotators ranked the model second of five, behind only Claude Opus 5 (vendor claim). It also says the model resists 93.3% of attacks on Lakera's B3 prompt-injection benchmark (vendor claim).
The first independent composite is more modest. On Artificial Analysis's Intelligence Index, which combines ten evaluations including Terminal-Bench 4.0, AutomationBench and Humanity's Last Exam, the preview scores 38. That places it 64th of 225 models in its comparison class and behind the strongest open-weight models.

Artificial Analysis figures checked on 7 October 2026, at list prices, with our estimate at the preview discount.
Three caveats belong next to that number. Mistral says the reinforcement learning run behind the preview is still in flight and expects "large and rapid improvements", so the score may move within weeks. Artificial Analysis currently lists the model as proprietary because the weights are not out yet. And LMArena, the crowd-voted leaderboard, has no entry for Mistral Large 4 on its text, vision or web development boards as of today.
So the honest summary for anyone searching for a Mistral Large 4 benchmark: Mistral's own numbers are strong on coding, security and business workflows, and the one independent composite available puts the preview in the second tier of open models. A full independent picture will take until the weights are out and more evaluators have run it.
Cheap per token can still be expensive per task
At list price, Mistral Large 4 costs about two-thirds of GPT-6.1 Sol's $2 input price and well under half of Sol's $10 output price. On a pricing page, that looks like a clear saving.
Artificial Analysis tells a different story. To run its full index, the Mistral preview generated 200 million output tokens, about three times the 67 million GPT-6.1 Sol needed. The result is a cost of $1.13 per index task for Mistral against $0.72 for Sol, with Sol scoring 52 to Mistral's 38. Even if the 50% preview discount applies across all token types, our estimate puts Mistral at roughly $0.57 per task, which is close to Sol and still well below it on score.

The per-token price wins on the pricing page; the token count decides the bill.
This is the same pattern we found when three models at an identical list price produced a tenfold spread in cost per task. Verbosity decides the bill. Mistral Large 4 is cheaper per task than GLM-5.3, Kimi K3 and Qwen3.8 on Artificial Analysis's numbers, so among large open models it is competitive on cost. Against the closed models most teams already run, the price argument is weaker than the per-token rate suggests.
Why the open weights and European operation are the real case
The strongest reason to look at Mistral Large 4 is control. Mistral says the preview runs on the same European infrastructure it was trained on, and that the model will be offered across regions including a European deployment Mistral operates "end-to-end, independently of other digital service providers and under European law." Once the weights are released, teams can also run the model on their own infrastructure.
For regulated teams, that combination matters. A bank, insurer or public body with data-residency rules can keep prompts and outputs inside a jurisdiction it chooses, pin a version that will not change under it, and keep running if a provider changes its terms. We argued in our piece on open-weight versus closed models that the enterprise decision usually comes down to this kind of control, and Mistral Large 4 is a clean test of that argument.
Mistral makes a related point about security work. It says several closed models score near zero on one vulnerability-reproduction test because they refuse the task, while its model can run "under their own policies" when self-deployed (vendor claim). Whether that fits your organisation is a governance question as much as a technical one.
The licence is the open question. Mistral Large 3 shipped under Apache 2.0. Mistral's documentation lists Large 4 as "Open" without naming a licence, and VentureBeat reports the weights will arrive on 27 October under a custom Mistral licence whose terms were not detailed. Until those terms are published, nobody can say whether commercial self-hosting will be unrestricted. As our Mistral model page puts it, European origin is one input to sovereignty. The region, the contract, the licence and the exit plan have to prove the rest.
How to evaluate Mistral Large 4 on your own agent steps
Launch benchmarks average across tasks your agents may never perform. The decision that matters is step by step: whether this model does a specific job in your workflow at least as well, at an acceptable cost, under the controls you need.
A practical sequence for the next month:
1. Pick candidate steps where control is the constraint. Multilingual document handling for EU entities, image-heavy extraction such as drawings or scanned forms, and security triage are the obvious ones given what Mistral emphasises.
2. Run your own evaluation set, with real inputs and known-good outputs, against the model currently on that step. Our guide to running evals on AI agents covers how to build one.
3. Measure cost per completed task, tokens included, at list price rather than the preview discount.
4. Wait for the licence and weights before committing any step that depends on self-hosting.
5. Re-test after the weights release, since Mistral expects the model to improve quickly.
This is easiest when the model is a setting on each step rather than a property of the whole agent, which is how assigning models per step works in Beam. Each step starts from a default drawn from the workspace's preferred models, can carry fallback models if a call fails, and can retry when accuracy drops below a threshold. Trying a new model on one step leaves the rest of the agent untouched. Beam runs models across several regions, including in Europe, so where a step runs is part of the same decision.
Common questions about Mistral Large 4
When is the Mistral Large 4 release date?
Mistral released Mistral Large 4 as a public preview on 6 October 2026, through its preview API on Mistral Studio. Mistral says the open weights will be released by the end of October 2026. VentureBeat reports a date of 27 October, although Mistral's own announcement only commits to the end of the month.
How does Mistral Large 4 do on benchmarks?
Mistral reports 61.7% on DeepSWE v1.1 and 59.9% on AutomationBench, both vendor claims. The first independent composite, Artificial Analysis's Intelligence Index, scores the preview at 38, behind GLM-5.3 at 45 and Kimi K3 at 44 but ahead of DeepSeek V4 Pro at 36. LMArena has not listed it yet.
How much does Mistral Large 4 cost?
The list price is $1.36 per million input tokens, $0.14 per million cached input tokens and $4.18 per million output tokens. During the preview, Mistral's model card shows half those rates: $0.68, $0.07 and $2.09. No end date for the discount is published. Artificial Analysis puts its cost at $1.13 per index task at list price.
Are Mistral Large 4's weights open, and under what licence?
Not yet. Mistral says it will release the weights by the end of October 2026, after red-teaming with security partners and state authorities. The licence has not been published. Mistral Large 3 used Apache 2.0, while VentureBeat reports Large 4 will use a custom Mistral licence. Check the terms before planning commercial self-hosting.
What does "le Chonk" mean?
It is Mistral's official nickname for Mistral Large 4, a nod to its size: about 1.05 trillion total parameters. Because it is a mixture-of-experts model, only around 49 to 52 billion of those parameters run for each token, which keeps it fast relative to its size but still demanding to host.





