Nvidia NeMo Switchyard Routes AI Agents to the Right Model

Share:
Audio Loading voice…
Nvidia NeMo Switchyard Routes AI Agents to the Right Model

Synopsis

Nvidia has introduced NeMo Switchyard, a tool that dynamically routes agent workflow steps across a model pool based on developer-defined quality, latency, and cost criteria. Executive Kari Briski explained the tool's rationale on MTS Live, highlighting the practical economics of multi-model AI pipelines.

Key Takeaways

Nvidia NeMo Switchyard routes individual agent workflow steps to different models based on quality, latency, and cost criteria set by developers.
Kari Briski , an Nvidia executive, presented the tool's rationale on MTS Live .
The tool addresses a core production challenge: not every pipeline step requires the same model, and forcing uniformity wastes resources.
Switchyard is part of Nvidia's broader expansion of the NeMo platform to support complex, multi-model AI agent systems.
Model routing allows teams to balance performance, speed, and inference cost dynamically rather than at design time.
Not every AI task deserves the same model — and Nvidia is building the infrastructure to make that distinction automatic. The chip giant unveiled NeMo Switchyard, a developer tool that routes each step of an agent workflow to the most suitable model from a chosen pool, weighing quality, latency, and cost in real time.
Speaking on MTS Live, Kari Briski, an Nvidia executive, laid out the core problem: agent workflows are not monolithic. A single pipeline might need a fast, cheap model for a classification step, a high-accuracy model for reasoning, and something in between for summarisation. Forcing every step through one model wastes money, slows pipelines, or sacrifices quality — sometimes all three at once.

Why model routing matters for production AI

NeMo Switchyard hands developers the controls. Rather than hard-coding a model into each workflow stage, developers define their own criteria — acceptable latency, quality thresholds, cost ceilings — and Switchyard dispatches each task accordingly. The result is a production deployment that can flex across a model pool instead of betting everything on one choice. This is a direct response to a real pain point. As AI agents move from demos into live products, the economics of inference become impossible to ignore. Running a frontier large language model on every micro-task in a pipeline is neither fast nor affordable at scale. Routing is how teams reconcile capability with cost.

NeMo's expanding role in the agentic era

Switchyard is part of Nvidia's broader push to make its NeMo platform the backbone of enterprise AI agent systems. The platform has steadily grown from a model-training toolkit into a full-stack environment for deploying multi-model, multi-step agent pipelines. Model routing is the connective tissue that makes those pipelines practical. For AI developers, the pitch is clear: stop treating model selection as a one-time architectural decision and start treating it as a dynamic, per-task optimisation. Switchyard is Nvidia's answer to what that looks like in code.

Point of View

Multi-model pipelines mirrors how cloud computing evolved from monolithic servers to microservices: granularity reduces waste. For Nvidia, owning the routing layer is strategically significant — it embeds the company deeper into the production stack, well beyond the GPU silicon. Developers who standardise on NeMo Switchyard effectively make Nvidia infrastructure a default dependency in their agent architectures.
NationPress
21 Aug 2026

Frequently Asked Questions

What is Nvidia NeMo Switchyard?
NeMo Switchyard is an Nvidia developer tool that automatically routes each step of an AI agent workflow to the most appropriate model in a pool, based on criteria like quality, latency, and cost that developers define themselves.
Why do AI agent workflows need model routing?
Different steps in an agent pipeline have different requirements — some need speed, others need accuracy, and some must minimise cost. Routing each step to the right model avoids the waste and bottlenecks of using a single model for everything.
Who is Kari Briski at Nvidia?
Kari Briski is an Nvidia executive who appeared on MTS Live to explain the case for model routing in AI agent workflows and the role of NeMo Switchyard.
What is the Nvidia NeMo platform?
NeMo is Nvidia's platform for building and deploying AI models and agent systems. It has expanded from a model-training toolkit into a full-stack environment supporting multi-model, multi-step agent pipelines.
How does NeMo Switchyard decide which model to use?
Developers set their own quality, latency, and cost thresholds; Switchyard then dispatches each workflow task to the model in the pool that best meets those criteria at runtime.
Nation Press
The Trail

Connected Dots

Tracing the thread behind this story — newest first.

8 Dots
  1. Latest 1 week ago
  2. 2 weeks ago
  3. 2 weeks ago
  4. 2 weeks ago
  5. 1 month ago
  6. 1 month ago
  7. 1 month ago
  8. 2 months ago
Google Prefer NP
On Google