Google Gemma logo
Google Gemma logo

Google Gemma

Google

MODEL DIRECTORY

MODEL DIRECTORY

Google Gemma

Google Gemma

Gemma 4 combines a practical size ladder, multimodality, broad language coverage, and Apache 2.0 weights. Google’s benchmark table is strong; independent evaluation of the 31B model is more moderate, so variant-specific testing matters.

CURRENT MODEL SNAPSHOT

Provider: Google
Current anchor: Gemma 4, with 31B as the reference variant
Lifecycle: Current
Weights: Open-weight under Apache 2.0
Reviewed: 21 July 2026

Abstract blue and purple gradient

Gemma 4 is an ecosystem play, not a single checkpoint

Google released Gemma 4 on 2 April 2026 as a family ranging from edge-oriented E2B and E4B variants through 12B, 26B-A4B, and 31B models. The larger variants support up to a 256,000-token context window, while smaller configurations target 128,000 tokens. Text, images, and video are supported across the family, with audio available in selected smaller or unified variants.

The family is pretrained across more than 140 languages, with over 35 advertised as ready out of the box. Apache 2.0 weights and the Google developer ecosystem make Gemma unusually approachable for teams that need model choice across device, private infrastructure, and managed endpoints.

That breadth is also a source of ambiguity. “Gemma 4” is not one performance number. A buyer must name the exact parameter count, active-parameter configuration, modality, context limit, quantization, serving runtime, and prompt format.

Google’s table is impressive; independent results are more restrained

For Gemma 4 31B, Google reports 85.2 on MMLU-Pro, 89.2 on AIME 2026, 80.0 on LiveCodeBench v6, 84.3 on GPQA Diamond, 76.9 on Tau2, and 76.9 on MMMU-Pro. The model card also reports 66.4 on MRCR at 128K. These are provider results and should be read with their task prompts, sampling settings, and comparison model versions.

Artificial Analysis reports an Intelligence Index score of 29 for Gemma 4 31B, around 35 output tokens per second, and a 256K window for the evaluated endpoint. That does not invalidate Google’s task scores; it shows that broad independent reasoning performance is more middle-of-pack than the launch table alone suggests.

  • Gemma’s strength is deployment breadth and model choice, not a claim to win every frontier benchmark.

  • Provider and independent suites answer different questions and must be dated and configuration-pinned.

  • Small variants can win on privacy, latency, and cost even when their aggregate intelligence score is lower.

The enterprise case: flexible, open, and operationally legible

Gemma is well suited to teams that need a defensible size ladder, multimodal document work, private deployment, or integration with Google’s tooling. It is less suitable when the requirement is simply the highest available general reasoning score regardless of cost or deployment control.

How we would evaluate it

Choose two variants that fit the real latency and infrastructure envelope. Run the same multilingual, multimodal workflow with citations, a strict schema, and adversarial evidence. Compare completion quality, unsupported claims, context retrieval, throughput, memory, energy, and reviewer time—not only benchmark averages.

Evidence used

Beam AI support status

Under evaluation. This record does not confirm support for every Gemma variant, runtime, modality, or region. Pin the actual artifact and serving path before rollout.

Use case 1

Private multimodal document agent

Run a selected Gemma variant on PDFs, tables, screenshots, and scanned evidence inside the approved environment. Score extraction accuracy, citations, unsupported inference, latency, and resource use.

Use case 2

Edge-to-cloud model routing

Route privacy-sensitive or latency-sensitive tasks to a small Gemma variant and complex cases to 31B or a premium model. Measure escalation quality, blended cost, completion, and review debt.

Use case 3

Multilingual knowledge workflow

Test the same source-linked process across priority languages, including mixed-language documents and ambiguous terminology. Measure accuracy, citation support, language drift, and reviewer correction.

Use case 4

Open-weight serving comparison

Benchmark two Gemma sizes across an approved runtime and managed endpoint. Include quantization, memory, throughput, safety layer, artifact provenance, patching, rollback, and total cost.

Related LLMs

Curated alternatives to compare before selecting a model family.

Start Today

Build AI agents with the right model

See how Beam can orchestrate governed AI workflows across the model family that fits your requirements.

Start Today

Build AI agents with the right model

See how Beam can orchestrate governed AI workflows across the model family that fits your requirements.

Start Today

Build AI agents with the right model

See how Beam can orchestrate governed AI workflows across the model family that fits your requirements.

FAQs

Frequently Asked Questions

Model selection, deployment, governance, and Beam support questions answered.

Is Gemma 4 one model?

No. It is a family with multiple sizes and modality configurations. Benchmark, context, hardware, and deployment claims must name the exact variant rather than using the family name alone.

How strong is Gemma 4 31B?

Google reports strong task benchmarks, including 85.2 on MMLU-Pro and 89.2 on AIME 2026. Artificial Analysis reports a broader Intelligence Index score of 29 for its evaluated 31B endpoint, showing a more moderate independent picture.

Can Gemma 4 be used commercially?

Google publishes Gemma 4 weights under Apache 2.0. Teams should still verify the exact repository, model card, notices, third-party dependencies, and acceptable-use requirements.

Which Gemma 4 size should an enterprise choose?

Start from latency, memory, privacy, modality, and quality requirements. Evaluate at least one small and one larger candidate on the same accepted-outcome test rather than assuming the largest model is automatically best.

Does Beam support Gemma 4 in production?

Beam support is currently Under evaluation. No blanket production-support claim applies across all Gemma sizes, runtimes, regions, and modalities.