Nvidia Launches NeMo Switchyard to Cut AI Computing Costs by Up to 74%
Enterprises deploying generative AI often rely on large models for every request, driving up computing costs and response times even when simpler systems could handle much of the workload. Nvidia’s NeMo Switchyard acts as an agent layer between applications and models, assessing task complexity and routing requests to an appropriate option. The approach is designed to balance accuracy, latency and cost as companies scale AI services.
Nvidia unveiled NeMo Switchyard alongside its Nemotron 3.5 model family, pairing the routing layer with a range of models suited to different workloads. Complex requests can be directed to larger, more capable models, while routine tasks are assigned to lower-cost small models. Nvidia said the dynamic routing strategy can reduce the total cost of completing AI tasks by as much as 74% without sacrificing accuracy.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.