Xiaomi MiMo
Xiaomi
Xiaomi MiMo-V2.5 pairs a 1M-token multimodal system with extremely low API pricing, while the Pro variant scales to 1T total parameters with 42B active. The strongest long-horizon claims remain provider-run and variant-specific.
CURRENT MODEL SNAPSHOT
Provider: Xiaomi MiMo Team
Current anchor: MiMo-V2.5 and V2.5-Pro
Lifecycle: Current
Weights: Open-weight under MIT terms
Reviewed: 21 July 2026

MiMo-V2.5 is Xiaomi’s full-stack bid for agentic AI
Xiaomi opened MiMo-V2.5 in public beta on 23 April 2026 and updated the family on 29 June. The standard V2.5 model is natively multimodal across text, image, video, and audio. MiMo-V2.5-Pro scales to one trillion total parameters while activating about 42 billion per token, with a one-million-token context window.
The distinction between variants matters. Provider demonstrations for the Pro model should not be attributed automatically to the standard V2.5 endpoint, and independent evaluations of V2.5 do not necessarily describe Pro. Enterprise buyers should record the exact identifier, reasoning mode, context configuration, and release date in every benchmark.
Xiaomi released weights under MIT terms, which makes the family unusually permissive for commercial experimentation and self-managed deployment. That legal openness does not make a trillion-parameter checkpoint operationally small.
The price story is exceptional; the benchmark story needs labels
Artificial Analysis reports the standard MiMo-V2.5 endpoint at an Intelligence Index score of 37, about 63 output tokens per second, and listed pricing near $0.14 per million input tokens and $0.28 per million output tokens. That is an aggressive cost profile, but it is not an independent score for V2.5-Pro.
Xiaomi reports that Pro completed a long coding task spanning 8,192 lines of code, 1,868 tool calls, and 11.5 hours. The company also reports 42% lower token use than Kimi K2.6 in one comparison and a 50% advantage over Muse Spark for the standard model. These results show the intended operating envelope; they need reproduction under a buyer’s own scaffold and acceptance tests.
Low token prices make broad workflow evaluation affordable.
Pro and standard results must remain separated in reporting.
Long task duration is not automatically a virtue; completion quality, recovery, and reviewer effort determine value.
The enterprise case: permissive weights and very low unit economics
MiMo deserves evaluation for high-volume multimodal extraction, long-running engineering agents, and organizations that want both hosted and self-managed options. The main risk is not lack of capability; it is assuming a provider demonstration transfers unchanged to another infrastructure and tool environment.
How we would evaluate it
Run the same task on standard V2.5 and Pro with fixed prompts, tools, time limits, and acceptance tests. Report completion, unsupported claims, tool failures recovered, total tokens, cost, latency, and reviewer time. For self-hosting, add quantization loss, throughput per accelerator, operational staffing, and license provenance.
Evidence used
Beam AI support status
Under evaluation. This page does not confirm a Beam integration, approved region, or production recommendation. Validate the exact MiMo variant and deployment surface before committing.
Use case 1
Long-running coding agent
Give standard V2.5 and Pro the same multi-repository change with tests and deliberate tool failures. Score accepted patches, regressions, recovery, token use, wall-clock time, and reviewer corrections.
Use case 2
Multimodal intake at scale
Process documents, screenshots, audio, and video into a typed case record with source links. Measure field accuracy, omissions, unsupported inferences, latency, and cost per accepted record.
Use case 3
Cost-sensitive agent routing
Route simple tasks to MiMo and difficult cases to a premium model. Measure escalation accuracy, blended completion rate, total token cost, and whether cheap first-pass work creates expensive review debt.
Use case 4
MIT-licensed deployment study
Compare the hosted service with a self-managed weight release. Include artifact provenance, quantization, accelerator needs, throughput, observability, safety controls, patching, and total operating cost.

Related LLMs
Curated alternatives to compare before selecting a model family.
FAQs
Frequently Asked Questions
Model selection, deployment, governance, and Beam support questions answered.
What is the difference between MiMo-V2.5 and V2.5-Pro?
V2.5 is the broadly evaluated multimodal release; Pro is the larger one-trillion-parameter, 42B-active variant used in Xiaomi’s most ambitious long-horizon demonstrations. Do not mix their benchmark or cost claims.
How strong is MiMo-V2.5 independently?
Artificial Analysis reports a score of 37 for the standard V2.5 endpoint, about 63 output tokens per second, and very low listed prices. That result should not be relabeled as a Pro score.
Are MiMo weights commercially usable?
Xiaomi states that the V2.5 weights are released under MIT terms, which is permissive for commercial use. Legal teams should still verify the exact repository, notices, dependencies, and artifact provenance.
Does a 1M context make MiMo suitable for every long task?
No. Long-context capacity must be tested for retrieval stability, instruction retention, tool discipline, latency, and cost. Many workflows still perform better with explicit retrieval and staged checkpoints.
Does Beam support Xiaomi MiMo in production?
Beam support is currently Under evaluation. No approved production integration or Beam benchmark is attached to this CMS record.





