Mark RadarMARK RADAR
About
EN
Sign in
Event File AI NVIDIA

NVIDIA Launches Nemotron 3.5 Lightning for AI Agents

1 reports · First detected 2026-08-12 · Last active 2026-08-12

Agentic AI systems repeatedly call tools, retain state and validate outputs, making inference speed and cost critical as workflows grow longer. NVIDIA is building an integrated stack of open Nemotron models, NeMo software and optimized DGX and RTX hardware to capture that demand. The strategy gives developers models they can deploy and customize while drawing more agent workloads onto NVIDIA’s computing platforms.

NVIDIA released Nemotron 3.5 Lightning on Aug. 11, 2026, with a mixture-of-experts architecture containing 30 billion parameters but activating 3 billion per token. It also open-sourced NeMo Switchyard, which routes requests to different models based on the task. In a CodeRabbit trial involving 1,000 routing tasks, post-training took less than three hours and cost under $100, while estimated serving costs for the workload fell to $1.16 from $2.34.

All Coverage

1 original reports

The Backstory

The history behind this event
Nvidia Open-Sources Nemotron 3.5 Lightning and Switchyard Router2026-08-12 · 2 reports · similarity 0.84

Nvidia introduced the Nemotron 3 family in December 2025 as an open foundation for agentic AI, where software must reason, call tools and coordinate multistep work over extended periods. Such systems can become slow and costly if every request is sent to a frontier model. A smaller, customizable execution model paired with a routing layer therefore matters for companies seeking to scale autonomous agents while balancing accuracy, latency and computing expense.

On Aug. 11, 2026, Nvidia open-sourced Nemotron 3.5 Lightning and NeMo Switchyard. Lightning is a 30-billion-parameter mixture-of-experts model that activates 3 billion parameters per token and targets frequent tool calls, state management and result validation. Nvidia said output speed can be as much as four times higher. Switchyard automatically directs each request to the most suitable model, aiming to cut processing time and operating costs. The company did not disclose pricing or an investment amount for either release.

Mark Radar|MARK RADAR
All times are in Taipei time (GMT+8)