MODEL DIRECTORY

MODEL DIRECTORY

MiniMax

MiniMax

MiniMax M3 is a price-performance outlier: a 428B-parameter open-weight MoE with only 23B active parameters, a 1M-token window, native multimodality, and unusually strong speed. Its community license and self-hosting demands still need careful review.

CURRENT MODEL SNAPSHOT

Provider: MiniMax
Current anchor: MiniMax M3
Lifecycle: Current
Weights: Open-weight under the MiniMax Community License
Reviewed: 21 July 2026

Abstract blue and purple gradient

MiniMax M3 turns sparse scale into a practical agent model

Released on 1 June 2026, MiniMax M3 is a 428-billion-parameter mixture-of-experts model that activates about 23 billion parameters per token. Its Multi-Scale Attention design is intended to keep a one-million-token context window usable without paying dense-model cost on every token. The same family handles text, images, and video, positioning M3 as an agent model rather than a chat-only release.

The architecture is significant because it targets the part of enterprise AI that benchmarks often miss: sustained work across tools, files, and long intermediate states. A large context window is useful only if retrieval remains stable, tool calls remain disciplined, and the model can recover from a bad intermediate step. M3’s sparse routing and high measured throughput make that evaluation economically plausible.

The weights are available, but “open-weight” is the accurate description. The MiniMax Community License is not equivalent to Apache 2.0 or MIT and includes commercial conditions that legal and procurement teams should review before product embedding or redistribution.

What the benchmarks say—and what they do not

Artificial Analysis currently reports an Intelligence Index score of 44, roughly 105–108 output tokens per second, a one-million-token context window, and listed API pricing around $0.30 per million input tokens and $1.20 per million output tokens. Earlier launch coverage cited a score of 55 under a previous evaluation methodology; those numbers are not directly comparable. The dated methodology matters more than the headline.

MiniMax reports 83.5 on BrowseComp and 37.1 on PostTrainBench. Its most interesting evidence is operational: in a 24-hour CUDA-optimization exercise, M3 produced 147 submissions across 1,959 tool calls and a reported 9.4× utilization improvement. That demonstrates long-horizon behavior under the provider’s scaffold; it does not prove the same result on another tool stack or hardware target.

  • The independent price and throughput profile makes M3 a serious challenger for high-volume agents.

  • Provider-run browser and coding tests should be reproduced with the buyer’s tools, timeouts, and acceptance tests.

  • A one-million-token limit is capacity, not evidence of perfect retrieval or instruction retention across that window.

The enterprise case: high leverage with license and serving work

M3 is attractive where a team needs long-context research, multimodal review, or repository-scale tool use at a much lower token price than premium closed models. It is less compelling when legal simplicity, a small serving footprint, or mature managed-service controls matter more than raw price-performance.

How we would evaluate it

Run one complete process with browser and API tools, contradictory evidence, a forced tool failure, a strict output schema, and an approval checkpoint. Measure accepted outcomes, unsupported claims, recovery rate, tail latency, total tokens, and reviewer time. For self-hosting, add serving throughput, quantization loss, observability, patching, and on-call effort.

Evidence used

Beam AI support status

Under evaluation. This page is a model-selection reference, not confirmation of a Beam integration or production recommendation. Validate the exact endpoint or weight release, license, region, retention, and workload before committing.

Use case 1

Long-horizon engineering agent

Give M3 a bounded repository change that requires search, implementation, tests, documentation, and recovery from a failed tool call. Score accepted patches, regressions, reviewer corrections, output tokens, and elapsed time.

Use case 2

Multimodal research workflow

Combine documents, screenshots, tables, and web evidence into a source-linked decision memo. Insert contradictions and irrelevant material, then score citation support, missed evidence, unsupported conclusions, and reviewer time.

Use case 3

High-volume tool orchestration

Compare M3 with a premium closed model on thousands of short and medium tool tasks. Track completion rate, retries, tail latency, total token cost, and whether lower unit pricing survives real failure handling.

Use case 4

Open-weight serving study

Benchmark an approved self-managed build against the hosted API. Include license review, artifact provenance, quantization, throughput, safety controls, monitoring, patching, rollback, and total operating cost.

Related LLMs

Curated alternatives to compare before selecting a model family.

Start Today

Build AI agents with the right model

See how Beam can orchestrate governed AI workflows across the model family that fits your requirements.

Start Today

Build AI agents with the right model

See how Beam can orchestrate governed AI workflows across the model family that fits your requirements.

Start Today

Build AI agents with the right model

See how Beam can orchestrate governed AI workflows across the model family that fits your requirements.

FAQs

Frequently Asked Questions

Model selection, deployment, governance, and Beam support questions answered.

Is MiniMax M3 open source?

M3 is open-weight, but its MiniMax Community License is not the same as a permissive OSI-style license such as Apache 2.0 or MIT. Review commercial-use, redistribution, and product-embedding terms before deployment.

How strong is MiniMax M3 independently?

Artificial Analysis currently reports an Intelligence Index score of 44, about 105–108 output tokens per second, and very low listed token pricing. Earlier launch scores used a different methodology and should not be compared directly.

Does the 1M-token window remove the need for retrieval?

No. It expands capacity but does not guarantee that evidence is found, weighted correctly, or retained across a long workflow. Retrieval, chunking, citations, and context hygiene still need task-specific testing.

Can an enterprise self-host MiniMax M3?

Yes in principle because weights are available, but the 428B-parameter system still requires serious inference engineering. Capacity, quantization, accelerator memory, observability, patching, and license compliance belong in the business case.

Does Beam support MiniMax M3 in production?

Beam support is currently Under evaluation. No approved production integration or Beam benchmark is attached to this record; validate the exact surface, controls, region, and operating model first.