Nvidia Open-Sources Nemotron 3.5 Lightning and NeMo Switchyard
AI agents repeatedly call tools, maintain state and validate results, making inference cost and latency critical as companies move from prototypes to persistent production systems. Nvidia’s Nemotron lineup is aimed at that execution layer, where many routine tasks do not require a costly frontier model. Pairing a smaller, customizable model with automated routing could help enterprises reserve more capable systems for complex work while running high-volume requests more efficiently.
Nvidia released Nemotron 3.5 Lightning and the open-source NeMo Switchyard router on Aug. 11, 2026. The mixture-of-experts model has 30 billion parameters but activates only 3 billion per token, delivering as much as four times faster output. Nvidia said it achieved 86% accuracy on PinchBench and completed 10,000 tasks 35% faster than Qwen3.6 35B at comparable accuracy. Switchyard directs each request to the most suitable configured model, helping reduce execution time and operating costs.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.