Kimi
Moonshot AI
Kimi K3 is the rare model launch that moved both benchmarks and markets. Its agentic results are genuinely strong, but the 2.8T-parameter system is not yet an easy or fully proven enterprise deployment.
CURRENT MODEL SNAPSHOT
Provider: Moonshot AI
Current anchor: Kimi K3
Lifecycle: Current
Weights: Open weights announced for 27 July 2026
Reviewed: 20 July 2026

Kimi K3 is the model that moved markets before its weights were released
Moonshot AI released Kimi K3 on 16 July 2026 as a 2.8-trillion-parameter mixture-of-experts model with native vision and a one-million-token context window. Only 16 of its 896 experts are active per token; Kimi Delta Attention and Attention Residuals are the architectural ideas intended to make that scale practical for long tasks.
The remarkable part was not just the technical specification. On 17 July, Z.ai fell roughly 28% and MiniMax about 16% as investors repriced the competitive value of Chinese AI developers. US technology stocks also fell: the Nasdaq lost 1.4% and Nvidia 2.2%. Kimi was a catalyst, not the sole cause—corporate earnings, oil prices, and geopolitical risk also affected the trading day, and technology shares partly recovered the following Monday.
The market interpretation was still rational. If a Chinese lab can approach closed-model performance, price its API aggressively, and then release its weights, scarcity and pricing power across the sector look less secure. That is an economic signal, not a benchmark.
As of this review on 20 July 2026, the full weights are scheduled for release on 27 July and are not yet available. K3 should therefore be described as an announced open-weight release, not an open-weight model that can already be deployed.
What the tests actually say
The independent results justify the attention, but not the mythology. Artificial Analysis scored K3 at 57 on its Intelligence Index. In its commercial-task harness, K3 reached a GDPval-AA v2 Elo of 1668—above GPT-5.5 and Claude Opus 4.8 in that test, behind Claude Fable 5. K3 also led AutomationBench-AA at 53% and placed second on AA-Briefcase.
Those are meaningful results for AI agents because they test task completion, not just short-form question answering. They also expose the trade-offs:
Artificial Analysis estimated about $0.94 per completed benchmark task. That is close to leading proprietary models, but materially above lower-cost challengers such as GLM and DeepSeek.
K3 generated roughly 132 million output tokens across the suite. Better task economics do not necessarily mean concise operation.
Accuracy rose from 33% on K2.6 to 46% on K3, while the measured hallucination rate worsened from 39% to 51%. Capability and reliability moved in different directions.
Moonshot’s own launch note says K3 still ranks below Claude Fable 5 and GPT-5.6 Sol overall.
Provider-run demonstrations add useful evidence: 24-hour kernel optimization, a 48-hour chip-design task, a MiniTriton compiler, and research workflows spanning many tools. Treat these as design evidence, not neutral benchmarks. A buyer should reproduce the structure of the work—not the provider’s demonstration prompt—on its own repository, documents, tools, and failure cases.
The enterprise case: impressive, operationally young
K3 is a serious candidate for long-horizon automation, repository-scale coding, and multimodal research. Its strongest argument is sustained tool use at frontier-adjacent quality—not the one-million-token number on its own.
The operational picture is less mature. Moonshot temporarily paused new subscriptions after demand exceeded capacity. The company recommends 64 or more accelerators for self-hosting, so the future open-weight release will improve control without making deployment simple or cheap. Hosted API pricing is $3 per million uncached input tokens and $15 per million output tokens, with cached input at $0.30.
How we would evaluate it
Run a complete business process with contradictory evidence, tool timeouts, a deliberately bad intermediate result, expected abstention, and a hard output schema. Track completed outcomes, unsupported claims, reviewer corrections, recovery rate, tail latency, and total tokens. Compare it with both a top closed model and a cheaper open-weight candidate.
Evidence used
Beam AI support status
Under evaluation. This is a model-selection reference, not confirmation of a Beam integration, benchmark result, regional availability, or production recommendation. Validate the exact K3 surface and version in the intended workflow before making a customer commitment.
Use case 1
Repository-scale engineering agent
Give K3 a real, bounded change that requires repository search, implementation, tests, documentation, and recovery from a failed tool call. Measure accepted patches, regressions, security issues, reviewer corrections, output tokens, and whether its long horizon beats a decomposed workflow.
Use case 2
Long-running operations benchmark
Test a multi-hour back-office workflow with browser or API tools, checkpoints, resumability, and human approval. AutomationBench makes this a promising hypothesis; only your completion rate, failure recovery, and cost per accepted case can validate it.
Use case 3
Multimodal evidence-room analysis
Use documents, tables, screenshots, and images to build a source-linked issue map. Insert contradictions and irrelevant material, then score citation support, missed evidence, unsupported conclusions, data minimization, and reviewer time.
Use case 4
Open-weight deployment study
After the weights ship, compare the hosted endpoint with an approved self-managed build. Include artefact provenance, quantization, 64-plus-accelerator capacity, throughput, safety layer, monitoring, patching, total cost, and operational staffing—not only model quality.

Related LLMs
Curated alternatives to compare before selecting a model family.
FAQs
Frequently Asked Questions
Model selection, deployment, governance, and Beam support questions answered.
Did Kimi K3 cause the July 2026 AI stock sell-off?
It was an important catalyst, especially for directly exposed Chinese AI developers: Z.ai fell roughly 28% and MiniMax about 16% on 17 July. The broader US sell-off also reflected earnings, oil, and geopolitical risk, and technology shares partly recovered the following Monday. The useful conclusion is that K3 challenged investor assumptions about frontier-model scarcity and pricing—not that one trading day proved it was the world’s best model.
How strong is Kimi K3 in independent benchmarks?
Artificial Analysis scored K3 at 57 on its Intelligence Index, 1668 Elo on GDPval-AA v2, and 53% on AutomationBench-AA, where it led at the time of review. The same evaluation found high output volume and a 51% hallucination rate, so the page treats the benchmark story as strong but not unqualified.
Is Kimi K3 open-weight today?
Not yet as of this review on 20 July 2026. Moonshot announced that the full weights would be released on 27 July. Until the artefacts ship and can be inspected, K3 should be described as a hosted model with an announced open-weight release.
Can an enterprise self-host Kimi K3 economically?
Possibly, but not casually. Moonshot recommends at least 64 accelerators, making capacity planning, inference engineering, safety controls, observability, patching, and on-call ownership central to the decision. Open weights reduce provider dependence; they do not remove infrastructure cost.
Does Beam support Kimi K3 in production?
Beam support is currently Under evaluation. No approved production integration or Beam benchmark is attached to this CMS record. Validate the endpoint or future weight release, region, data controls, tools, failure behaviour, and operating model before making a customer commitment.





