NVIDIA Launches NeMo Switchyard to Reduce AI Agent Operating Costs Using AI Model Routing Technology
Odaily Planet Daily reports that NVIDIA has launched the AI model routing system NeMo Switchyard, aimed at helping developers dynamically allocate AI agent tasks among multiple large models, ensuring performance while reducing costs and latency.
It states that building AI agents does not mean relying on a single large model. Different models have their advantages in inference capability, response speed, and operating costs. For instance, lightweight models may be suitable for classification tasks, while complex reasoning requires more powerful cutting-edge models. Sending all requests to the largest model will lead to increased costs and latency; conversely, using only smaller models may compromise the quality of complex task completion. NeMo Switchyard employs an intelligent routing mechanism to automatically select the most suitable execution model among various dedicated and cutting-edge models based on factors such as task requirements, model capabilities, costs, latency, and system status. This system allows developers to switch between different model vendors and versions without altering the application architecture.
It introduces multiple routing strategies provided by NeMo Switchyard, including a no-training-required LLM classifier routing, stage routing, escalation routing, and an adjustable routing model trained on real workload data. Among these, escalation routing prioritizes assigning tasks to lower-cost models and upgrades requests to stronger models when it detects increased task complexity, persistent errors, or execution stalls, thus achieving a balance between performance and cost.
It indicates that in relevant tests, NeMo Switchyard significantly reduced AI agent operating costs while maintaining a high task completion rate by distributing tasks among different models. For example, compared to using only cutting-edge models, the escalation routing-based solution reduced costs by approximately 74% in the LangChain multi-round agent test, with only 7% of requests needing to call on cutting-edge models.
Additionally, it has collaborated with companies such as Cognition, Nous Research, Ramp, LangChain, LiteLLM, and Kong to integrate NeMo Switchyard into AI agent development processes and enterprise application infrastructures.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

DDR4 Price Increase Expected to Continue into Q4, Maintaining 'Attractive' View on Greater China Semiconductor Industry

Anthropic and OpenAI May Limit API Access, Raising Competition Risks for Businesses

AI Race: The United States Demands to Choose a Side - The French Mistral Turns to China

What Were Meme Coins Selling? — A Reconfiguration of Belief | HashHub Research

Foreign Media: Underestimated Strong Performance of European Stock Markets Emerges

US ESG Funds See $3 Billion Net Inflow in Q2, Led by AI Power Grid ETFs

74% of Traditional Finance Investors Shift to Cryptocurrency Exchanges as 24-Hour Trading Changes Investment Patterns

Mitchell Green Points Out That the AI Market Is in a Bubble

Bitcoin Coalition Calls for Access to the Same AI Tools Used by Attackers

Microsoft Retreats from China: AI-Driven Strategy of 'Reduction Instead of Withdrawal' - Reuters

Hassabis Advocates for AI Ethics Committee and Regulatory Framework

IBM Signs Strategic Partnership Agreement with OpenAI

Korea Investment Holdings Selected as Preferred Negotiator for KDB Life

From Subsidy Narrative to Real Returns: The Turning Point Year for DePIN Power Networks

Kalshi and DoubleZero Launch Real-Time Market Data Subscription Service

Thrive Holdings Completes $2 Billion Financing, Valuation Reaches $12 Billion

AI Hyperscalers Are Pricing Bitcoin Miners Off The Grid— Here's Why Its A Massive Win-Win

Global Cybersecurity Alliance (GCSA) and Academy of Engineering and Technology of the Developing World (AETDEW) Jointly Establish AI Cybersecurity Research Center

Gold Price and Inflation: Can the US CPI Push Gold to $5,000?

Mianbi Intelligence Launches A-Share IPO Counseling, Raising Over 5 Billion Yuan in Six Months

AI Agents Expose Attack Capabilities, Cybersecurity Spending May Accelerate

Solana Foundation Supports Diversity in Perpetual Contract Ecosystem

Anthropic Plans IPO Meeting Ahead of September or Early October Launch

OpenAI Launches GPT-5.6-Cyber Model to Enhance Cybersecurity Task Capabilities

Over 40 Organizations Call for AI Labs to Grant Model Access to Bitcoin Defenders

Zhipu's User Base Approaches 7 Million, Activates Over 50,000 Domestic AI Chips

Vitalik Claims Minimax H3 Surpasses HunyuanVideo 1.5

DeepMind Leadership Change: Why Google Cloud Could Be the Biggest Winner?

Jefferies Lowers SanDisk Target Price to $1,750, Maintains Buy Rating








