NVIDIA Nemotron logo
NVIDIA Nemotron logo

NVIDIA Nemotron

NVIDIA

MODEL DIRECTORY

MODEL DIRECTORY

NVIDIA Nemotron

NVIDIA Nemotron

Nemotron 3 Ultra is a 550B/55B-active open-weight hybrid Mamba-Attention MoE built for long-running agents. Its throughput and transparency are exceptional; the 352GB NVFP4 checkpoint and NVIDIA-heavy serving path are not lightweight.

CURRENT MODEL SNAPSHOT

Provider: NVIDIA
Current anchor: Nemotron 3 Ultra
Lifecycle: Current
Weights: Open-weight with published data and recipes
Reviewed: 21 July 2026

Abstract blue and purple gradient

Nemotron 3 Ultra is designed around the economics of long-running agents

NVIDIA released Nemotron 3 Ultra on 4 June 2026. The model has 550 billion total parameters with about 55 billion active per token, a hybrid Mamba-Attention mixture-of-experts architecture, native NVFP4 support, and a one-million-token context window. NVIDIA publishes weights, training data information, and recipes, making the family unusually transparent for a model at this scale.

The architecture is tuned for sustained inference throughput rather than maximum dense attention at every step. That matters for agents that generate large traces, call many tools, or maintain long working memory. It also aligns closely with NVIDIA’s hardware and software stack, which can be an advantage for an existing NVIDIA estate and a dependency for everyone else.

The NVFP4 checkpoint is approximately 352GB. Open weights improve control and auditability; they do not make Ultra a small deployment.

The strongest evidence is throughput plus agent-task coverage

At release, Artificial Analysis reported an Intelligence Index score of 47.7 under its then-current methodology and throughput above 400 output tokens per second on a Blackbox-hosted configuration. Current model pages may show a different normalized score as the suite evolves, so reports should preserve the evaluation date and methodology.

NVIDIA reports 91 on PinchBench, 33 on EnterpriseOps-Gym, 54 on Terminal-Bench 2.0, 82 on IFBench, 1,448 on GDPval-AA, 56 on ProfBench Search, and 95 on RULER at one million tokens. The company also reports substantial throughput advantages over selected open peers and about 30% lower task cost in its evaluation.

  • Independent release testing supports the claim that Ultra can be extremely fast on optimized infrastructure.

  • Provider benchmark tables still require exact hardware, runtime, sampling, and comparison-version context.

  • The business case should include accelerator availability and platform engineering, not only token price.

The enterprise case: transparent and powerful, with a large systems footprint

Nemotron is compelling for long-running agent orchestration, coding, and high-throughput private inference—especially where NVIDIA infrastructure and expertise already exist. It is a poor fit for teams that want a compact checkpoint, vendor-neutral serving, or minimal inference engineering.

How we would evaluate it

Run a multi-hour agent process with tool failures, checkpoints, and strict acceptance tests on the intended hardware. Measure accepted outcomes, recovery, tokens, throughput, time to first token, energy, accelerator occupancy, and reviewer effort. Compare the optimized NVIDIA path with one viable alternative, including staff cost and lock-in.

Evidence used

Beam AI support status

Under evaluation. No Beam benchmark or approved production topology is attached to this record. Validate the exact checkpoint, runtime, and infrastructure.

Use case 1

Long-running operations agent

Run a multi-hour process with APIs, checkpoints, resumability, approval, and deliberate failures. Score completed outcomes, recovery, unsupported actions, total tokens, accelerator occupancy, and reviewer time.

Use case 2

High-throughput coding fleet

Benchmark parallel repository tasks on the intended NVIDIA stack. Track accepted patches, regressions, throughput, queue time, energy, hardware utilization, and cost per accepted change.

Use case 3

Private one-million-token research

Use controlled internal evidence across a long context, then test retrieval with planted contradictions and distant dependencies. Score citations, omissions, unsupported conclusions, latency, and memory cost.

Use case 4

Open-model platform evaluation

Compare the full NVIDIA serving path with a vendor-neutral alternative. Include weights, runtime, quantization, observability, safety, patching, staffing, hardware supply, rollback, and three-year operating cost.

Related LLMs

Curated alternatives to compare before selecting a model family.

Start Today

Build AI agents with the right model

See how Beam can orchestrate governed AI workflows across the model family that fits your requirements.

Start Today

Build AI agents with the right model

See how Beam can orchestrate governed AI workflows across the model family that fits your requirements.

Start Today

Build AI agents with the right model

See how Beam can orchestrate governed AI workflows across the model family that fits your requirements.

FAQs

Frequently Asked Questions

Model selection, deployment, governance, and Beam support questions answered.

What is Nemotron 3 Ultra?

It is NVIDIA’s 550B-parameter, 55B-active hybrid Mamba-Attention MoE for long-running reasoning and agents, with a one-million-token context window and open weights, data information, and recipes.

How fast is Nemotron 3 Ultra?

Artificial Analysis reported more than 400 output tokens per second on an optimized Blackbox configuration at release. Throughput is highly dependent on hardware, runtime, batching, and quantization, so reproduce it on the intended stack.

How strong are its benchmarks?

NVIDIA reports strong agent and long-context results such as 91 on PinchBench and 95 on RULER at 1M. Artificial Analysis reported a 47.7 Intelligence Index at release under its then-current methodology. Preserve dates and settings when comparing.

Is Nemotron easy to self-host?

It is open-weight but not lightweight. The NVFP4 checkpoint is about 352GB, and obtaining the headline throughput requires serious accelerator capacity and NVIDIA inference expertise.

Does Beam support Nemotron 3 Ultra in production?

Beam support is currently Under evaluation. No approved production topology or Beam benchmark is attached to this record.