Google Gemma
Gemma 4 combines a practical size ladder, multimodality, broad language coverage, and Apache 2.0 weights. Google’s benchmark table is strong; independent evaluation of the 31B model is more moderate, so variant-specific testing matters.
CURRENT MODEL SNAPSHOT
Provider: Google
Current anchor: Gemma 4, with 31B as the reference variant
Lifecycle: Current
Weights: Open-weight under Apache 2.0
Reviewed: 21 July 2026

Gemma 4 is an ecosystem play, not a single checkpoint
Google released Gemma 4 on 2 April 2026 as a family ranging from edge-oriented E2B and E4B variants through 12B, 26B-A4B, and 31B models. The larger variants support up to a 256,000-token context window, while smaller configurations target 128,000 tokens. Text, images, and video are supported across the family, with audio available in selected smaller or unified variants.
The family is pretrained across more than 140 languages, with over 35 advertised as ready out of the box. Apache 2.0 weights and the Google developer ecosystem make Gemma unusually approachable for teams that need model choice across device, private infrastructure, and managed endpoints.
That breadth is also a source of ambiguity. “Gemma 4” is not one performance number. A buyer must name the exact parameter count, active-parameter configuration, modality, context limit, quantization, serving runtime, and prompt format.
Google’s table is impressive; independent results are more restrained
For Gemma 4 31B, Google reports 85.2 on MMLU-Pro, 89.2 on AIME 2026, 80.0 on LiveCodeBench v6, 84.3 on GPQA Diamond, 76.9 on Tau2, and 76.9 on MMMU-Pro. The model card also reports 66.4 on MRCR at 128K. These are provider results and should be read with their task prompts, sampling settings, and comparison model versions.
Artificial Analysis reports an Intelligence Index score of 29 for Gemma 4 31B, around 35 output tokens per second, and a 256K window for the evaluated endpoint. That does not invalidate Google’s task scores; it shows that broad independent reasoning performance is more middle-of-pack than the launch table alone suggests.
Gemma’s strength is deployment breadth and model choice, not a claim to win every frontier benchmark.
Provider and independent suites answer different questions and must be dated and configuration-pinned.
Small variants can win on privacy, latency, and cost even when their aggregate intelligence score is lower.
The enterprise case: flexible, open, and operationally legible
Gemma is well suited to teams that need a defensible size ladder, multimodal document work, private deployment, or integration with Google’s tooling. It is less suitable when the requirement is simply the highest available general reasoning score regardless of cost or deployment control.
How we would evaluate it
Choose two variants that fit the real latency and infrastructure envelope. Run the same multilingual, multimodal workflow with citations, a strict schema, and adversarial evidence. Compare completion quality, unsupported claims, context retrieval, throughput, memory, energy, and reviewer time—not only benchmark averages.
Evidence used
Beam AI support status
Under evaluation. This record does not confirm support for every Gemma variant, runtime, modality, or region. Pin the actual artifact and serving path before rollout.
Use case 1
Private multimodal document agent
Run a selected Gemma variant on PDFs, tables, screenshots, and scanned evidence inside the approved environment. Score extraction accuracy, citations, unsupported inference, latency, and resource use.
Use case 2
Edge-to-cloud model routing
Route privacy-sensitive or latency-sensitive tasks to a small Gemma variant and complex cases to 31B or a premium model. Measure escalation quality, blended cost, completion, and review debt.
Use case 3
Multilingual knowledge workflow
Test the same source-linked process across priority languages, including mixed-language documents and ambiguous terminology. Measure accuracy, citation support, language drift, and reviewer correction.
Use case 4
Open-weight serving comparison
Benchmark two Gemma sizes across an approved runtime and managed endpoint. Include quantization, memory, throughput, safety layer, artifact provenance, patching, rollback, and total cost.

Related LLMs
Curated alternatives to compare before selecting a model family.
FAQs
Frequently Asked Questions
Model selection, deployment, governance, and Beam support questions answered.
Is Gemma 4 one model?
No. It is a family with multiple sizes and modality configurations. Benchmark, context, hardware, and deployment claims must name the exact variant rather than using the family name alone.
How strong is Gemma 4 31B?
Google reports strong task benchmarks, including 85.2 on MMLU-Pro and 89.2 on AIME 2026. Artificial Analysis reports a broader Intelligence Index score of 29 for its evaluated 31B endpoint, showing a more moderate independent picture.
Can Gemma 4 be used commercially?
Google publishes Gemma 4 weights under Apache 2.0. Teams should still verify the exact repository, model card, notices, third-party dependencies, and acceptable-use requirements.
Which Gemma 4 size should an enterprise choose?
Start from latency, memory, privacy, modality, and quality requirements. Evaluate at least one small and one larger candidate on the same accepted-outcome test rather than assuming the largest model is automatically best.
Does Beam support Gemma 4 in production?
Beam support is currently Under evaluation. No blanket production-support claim applies across all Gemma sizes, runtimes, regions, and modalities.






