HyperAIHyperAI

Command Palette

Search for a command to run...

17 hours ago
NVIDIA
Agent
LLM

NVIDIA Unveils Faster Nemotron 3.5 Lightning, NeMo Switchyard for Agentic AI

NVIDIA has expanded its open artificial intelligence ecosystem with the release of Nemotron 3.5 Lightning and NeMo Switchyard, targeting the growing demand for efficient, autonomous agent architectures. Launched on August 11, these tools address enterprise and developer needs for greater control over AI deployment, cost management, and specialized task execution. Nemotron 3.5 Lightning is a fully customizable, 30-billion-parameter mixture-of-experts open model engineered for high-volume, always-on agentic workloads. Developed with contributions from the Nemotron Coalition, the model achieves up to four times faster token generation and completes agentic tasks thirty percent faster than comparable open alternatives while maintaining frontier-level accuracy. Its open-weight architecture allows organizations to fine-tune the model on proprietary datasets using NVIDIA NeMo, optimizing it for domain-specific applications. Early deployments include CrowdStrike for cybersecurity monitoring, Harvey for legal research, CodeRabbit for automated code review, and Lila Sciences for life sciences reasoning. The model supports flexible deployment across the hardware continuum, running locally on NVIDIA RTX PCs, DGX Spark, and Jetson devices, while scaling seamlessly to data centers and cloud environments. Compatibility extends to leading inference frameworks, including vLLM, Ollama, llama.cpp, LM Studio, and Unsloth, ensuring streamlined integration for developers. Complementing the model is NeMo Switchyard, an open-source routing library that automates intelligent model selection within multi-agent workflows. Rather than relying on a single default model, enterprises can use Switchyard to dynamically direct prompts to the most appropriate model for each step based on predefined priorities such as accuracy, latency, and cost. Internal benchmarks indicate that this routing approach maintains frontier-level performance while reducing task completion costs to approximately one-third of those associated with standalone Opus 4.8 deployments. The library allows developers to implement custom routing algorithms without rewriting existing applications, significantly lowering integration overhead and improving overall token economics. The releases underscore NVIDIA’s strategic pivot toward supporting decentralized, efficient AI agent systems. By providing transparent training methodologies and open weights, NVIDIA enables enterprises to maintain data privacy, conduct thorough audits, and build proprietary reasoning capabilities. Both Nemotron 3.5 Lightning and NeMo Switchyard are available through Hugging Face, ModelScope, OpenRouter, and build.nvidia.com, with Switchyard also accessible via GitHub. The launch aligns with NVIDIA’s broader August initiative to accelerate the local AI ecosystem, reinforcing industry momentum toward autonomous, cost-effective, and highly customizable AI deployments.

Related Links