Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status

Nemotron 3.5 Lightning

Budget

by NVIDIA

Nemotron 3.5 Lightning is NVIDIA's August 2026 open model, and it is shipped as a starting point rather than a finished assistant. The architecture is a hybrid of Mamba-2 and mixture of experts at 30 billion total parameters with 3 billion active, which is what makes the $0.08 input and $0.20 output pricing possible. NVIDIA reports 81.94 on MMLU Pro, 75.44 on GPQA Diamond, 51.56 on SWE-bench Verified, and 36.97 on BrowseComp. One correction worth carrying: the license is OpenMDW-1.1, an external open-model license, not an NVIDIA-authored one, which makes the terms more predictable than the Nemotron name might suggest. The card advertises up to 1M context but notes that roughly 256K is the practical ceiling on a single H100. NVIDIA positions this explicitly as a customization and post-training base for agentic workloads, so read the benchmark numbers as a floor you are expected to build on.

Input Price

$0.08

per 1M tokens

Output Price

$0.20

per 1M tokens

Context Window

262K

tokens

Released

2026-08

Open source

Capabilities

texttool-usecodereasoning

Key Strengths

  • $0.08/$0.20 per 1M tokens, among the cheapest models tracked
  • 30B total with only 3B active per token
  • OpenMDW-1.1, a standard external open-model license
  • Hybrid Mamba-2 and MoE architecture
  • Built and licensed as a post-training and customization base

Best For

  • Fine-tuning and post-training for agentic tasks
  • High-volume classification and extraction
  • Latency-sensitive tool-calling loops
  • Edge and single-GPU deployments

Pricing Details

Input tokens

$0.08

per 1M tokens

Output tokens

$0.20

per 1M tokens

Estimated cost per 1K requests

$0.18

~1K input + ~500 output tokens avg

Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.

Open Source Model

Nemotron 3.5 Lightning is free to download and self-host under the OpenMDW-1.1. Hosted API pricing varies by provider (e.g., Together, Fireworks, Groq). See our open source LLM guide for deployment options.

Related Models

View DocumentationCompare ModelsCost CalculatorFull Pricing Guide