NVIDIA Nemotron
NVIDIA
Nemotron 3 Ultra is a 550B/55B-active open-weight hybrid Mamba-Attention MoE built for long-running agents. Its throughput and transparency are exceptional; the 352GB NVFP4 checkpoint and NVIDIA-heavy serving path are not lightweight.
CURRENT MODEL SNAPSHOT
Provider: NVIDIA
Current anchor: Nemotron 3 Ultra
Lifecycle: Current
Weights: Open-weight with published data and recipes
Reviewed: 21 July 2026

Nemotron 3 Ultra is designed around the economics of long-running agents
NVIDIA released Nemotron 3 Ultra on 4 June 2026. The model has 550 billion total parameters with about 55 billion active per token, a hybrid Mamba-Attention mixture-of-experts architecture, native NVFP4 support, and a one-million-token context window. NVIDIA publishes weights, training data information, and recipes, making the family unusually transparent for a model at this scale.
The architecture is tuned for sustained inference throughput rather than maximum dense attention at every step. That matters for agents that generate large traces, call many tools, or maintain long working memory. It also aligns closely with NVIDIA’s hardware and software stack, which can be an advantage for an existing NVIDIA estate and a dependency for everyone else.
The NVFP4 checkpoint is approximately 352GB. Open weights improve control and auditability; they do not make Ultra a small deployment.
The strongest evidence is throughput plus agent-task coverage
At release, Artificial Analysis reported an Intelligence Index score of 47.7 under its then-current methodology and throughput above 400 output tokens per second on a Blackbox-hosted configuration. Current model pages may show a different normalized score as the suite evolves, so reports should preserve the evaluation date and methodology.
NVIDIA reports 91 on PinchBench, 33 on EnterpriseOps-Gym, 54 on Terminal-Bench 2.0, 82 on IFBench, 1,448 on GDPval-AA, 56 on ProfBench Search, and 95 on RULER at one million tokens. The company also reports substantial throughput advantages over selected open peers and about 30% lower task cost in its evaluation.
Independent release testing supports the claim that Ultra can be extremely fast on optimized infrastructure.
Provider benchmark tables still require exact hardware, runtime, sampling, and comparison-version context.
The business case should include accelerator availability and platform engineering, not only token price.
The enterprise case: transparent and powerful, with a large systems footprint
Nemotron is compelling for long-running agent orchestration, coding, and high-throughput private inference—especially where NVIDIA infrastructure and expertise already exist. It is a poor fit for teams that want a compact checkpoint, vendor-neutral serving, or minimal inference engineering.
How we would evaluate it
Run a multi-hour agent process with tool failures, checkpoints, and strict acceptance tests on the intended hardware. Measure accepted outcomes, recovery, tokens, throughput, time to first token, energy, accelerator occupancy, and reviewer effort. Compare the optimized NVIDIA path with one viable alternative, including staff cost and lock-in.
Evidence used
Beam AI support status
Under evaluation. No Beam benchmark or approved production topology is attached to this record. Validate the exact checkpoint, runtime, and infrastructure.
Use case 1
Long-running operations agent
Run a multi-hour process with APIs, checkpoints, resumability, approval, and deliberate failures. Score completed outcomes, recovery, unsupported actions, total tokens, accelerator occupancy, and reviewer time.
Use case 2
High-throughput coding fleet
Benchmark parallel repository tasks on the intended NVIDIA stack. Track accepted patches, regressions, throughput, queue time, energy, hardware utilization, and cost per accepted change.
Use case 3
Private one-million-token research
Use controlled internal evidence across a long context, then test retrieval with planted contradictions and distant dependencies. Score citations, omissions, unsupported conclusions, latency, and memory cost.
Use case 4
Open-model platform evaluation
Compare the full NVIDIA serving path with a vendor-neutral alternative. Include weights, runtime, quantization, observability, safety, patching, staffing, hardware supply, rollback, and three-year operating cost.

Related LLMs
Curated alternatives to compare before selecting a model family.
FAQs
Frequently Asked Questions
Model selection, deployment, governance, and Beam support questions answered.
What is Nemotron 3 Ultra?
It is NVIDIA’s 550B-parameter, 55B-active hybrid Mamba-Attention MoE for long-running reasoning and agents, with a one-million-token context window and open weights, data information, and recipes.
How fast is Nemotron 3 Ultra?
Artificial Analysis reported more than 400 output tokens per second on an optimized Blackbox configuration at release. Throughput is highly dependent on hardware, runtime, batching, and quantization, so reproduce it on the intended stack.
How strong are its benchmarks?
NVIDIA reports strong agent and long-context results such as 91 on PinchBench and 95 on RULER at 1M. Artificial Analysis reported a 47.7 Intelligence Index at release under its then-current methodology. Preserve dates and settings when comparing.
Is Nemotron easy to self-host?
It is open-weight but not lightweight. The NVFP4 checkpoint is about 352GB, and obtaining the headline throughput requires serious accelerator capacity and NVIDIA inference expertise.
Does Beam support Nemotron 3 Ultra in production?
Beam support is currently Under evaluation. No approved production topology or Beam benchmark is attached to this record.





