UPDATED 09:00 EDT / AUGUST 11 2026

AI

Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard to give enterprise AI capability options

Artificial intelligence silicon and software giant Nvidia Corp. today announced two new services: a highly customizable Nemotron model and an agentic AI model router named NeMo Switchyard.

As enterprises find themselves drowning in artificial intelligence model options, the question is no longer raw power and capability, but fit-for-what-purpose and when. As agents become the norm, the tasks they perform range from sequences of swift, simple tasks to high-complexity reasoning across deep knowledge domains. The former might require small models that can swiftly handle small sorting tasks efficiently and the latter would be best handled by frontier models with high intelligence and reasoning capability.

Nemotron 3.5 Lightning joins the open model family to support high-volume tasks when powering always-on agents. It’s a 30 billion-parameter mixture-of-experts model that Nvidia claims is capable of delivering up to four times the output speed and 30% faster agentic task completion compared with other models in its weight class. It’s also open and customizable, meaning it can be readily post-trained with Nvidia NeMo on an enterprise’s own hardware, with proprietary domain data, tools and workflows to improve accuracy for specialized tasks.

This alone makes the model a powerful backbone workhorse within the family.

Nvidia launched the Nemotron 3 model family in December to act as an open foundation for agentic AI systems. At the time, it consisted of three sizes: Nano, Super and Ultra. The company said Lightning’s differentiator is that it can be specialized quickly.

A company can reproducibly regenerate its own data and specialization in a newly trained model with minimal equipment and cost.

“What we’re hearing is that Lightning is remarkably easy to customize,” said Vice President of Generative AI Kari Briski.

CodeRabbit Inc., according to Briski, used Nvidia’s standard auto model recipe, trained for one epoch, and produced a router agent for $85 in around two hours. She also said another partner dropped Lightning directly into an existing post-training stack “with no changes required.”

In one case she said, a company did training on a single H100 card relatively inexpensively, and another simply set up a training job overnight and came back in the morning to pick up the results. This is turning post-training and fine-tuning experience into loading the dishwasher and hitting the button.

Nvidia is also releasing its datasets, its post-training datasets and the recipes/framework specifically used to train the models. Developers could blend Nvidia’s own post-training data with their own enterprise data to create powerful hybrids that get the best of both worlds.

“That’s why we see people getting really great results really quickly,” Briski said.

Picking the right model for the task, why AI agents need a router

Cost, accuracy and efficiency aren’t “cheap model versus expensive model”; it’s the best fit for the task. When an enterprise agent is doing work, the tasks vary in complexity, expertise, token use, prediction and capability.  That means routing to a model that will do the task correctly is similar to providing a task to a person. Some models are better at coding, others are better at inferring context from text, and yet others have specialized knowledge instilled in them.

NeMo Switchyard is an open-source model routing library for AI agents that takes into account what numerous AI models are available and their capabilities, and routes prompts to the most capable and efficient model available at each step of the agent workflow based on specific needs.

“Depending on your routing strategy, it wants to choose the best model,” Briski explained.

Developers can customize the router with their own choice of strategy depending on their priority. For example, quality, delay and cost requirements. A system of models could be sorted for tokenomics, speed or particular quirks, expertise or other needs that an industry requires.

Nvidia said it’s working with a number of partners across the AI ecosystem on intelligent AI model routing alongside tools developers already use, including Boomi LP, Cadence Design Systems Inc., Classmethod Inc., Cognition AI Inc., Kong Inc., Langchain Inc., Nous Research Inc. and Siemens AG.

Image: Nvidia

A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.

Are you AWS customer?  Support SiliconANGLE Financially by buying your AWS services from our Marketplace portal page and links.  

About SiliconANGLE Media
SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.