Artificial intelligence is entering a new era. Rather than simply responding to prompts, AI is evolving into autonomous agents capable of planning, reasoning, using tools, and completing complex, multi-step workflows with minimal human intervention. From managing schedules and summarizing documents to analyzing business data and automating repetitive tasks, AI agents are poised to become valuable digital collaborators for everyone, including enterprises, organizations, and individuals.
As these agents become more capable, organizations need computing platforms that can support AI workloads closer to where data is generated and decisions are made. Solutions like the Acer Veriton GN100 AI Mini Workstation, powered by the NVIDIA GB10 Grace Blackwell Superchip, provide the local AI infrastructure needed to bring advanced AI capabilities closer to users while maintaining greater control over data and deployment.
As AI agents continue to evolve, however, one question becomes increasingly important: where should AI workloads ultimately be run?
One Agent, Many Models
Not every step in an AI workflow requires the largest or most powerful language model. This is where NVIDIA Nemotron 3.5 Lightning comes into play.
Designed as a customizable, high-throughput open model for powering always-on AI agents, Nemotron 3.5 Lightning is a 30B Mixture-of-Experts (MoE) model with 3B active parameters, distilled from NVIDIA’s frontier Nemotron 3 Ultra model. Built for local AI systems like the Acer Veriton GN100 AI Mini Workstation, Nemotron 3.5 Lightning brings agentic capabilities closer to users while providing the flexibility to customize and optimize models for specialized workflows.
Trained for popular agent harnesses, Nemotron 3.5 Lightning delivers strong out-of-the-box accuracy for agentic tasks and can be post-trained to improve performance for specific enterprise applications. With fast token generation and efficient token rollout, it helps always-on agents complete tasks faster across long-running, multi-turn workflows.
As an example, imagine an AI agent reviewing a financial service contract. It might begin by extracting text from a document, identifying key clauses, comparing them against company policies, flagging risks, retrieving supporting information from an internal knowledge base, and finally generating a summary with recommended actions.
Similarly, in cybersecurity operations, AI agents can assist security teams by analyzing alerts, classifying incidents, querying logs, correlating indicators, and preparing structured findings for analysts.
While some of these steps can be handled efficiently by specialized always-on, 24/7 models like Nemotron 3.5 Lightning, others may require the deeper reasoning capabilities of larger frontier models. Rather than relying on a single model for every stage, modern AI agents are increasingly built around systems of models - selecting the most appropriate model for each workflow step.
This approach allows organizations to balance speed, efficiency, and capability, using local AI models where they provide the most value while extending to larger models when additional intelligence is needed.
Extending Beyond Local with Hybrid AI
While local AI has advanced significantly, not every workload is best handled locally. Some tasks may require the advanced reasoning capabilities, broader knowledge, or specialized capabilities of larger frontier models.
This is where hybrid AI becomes essential.
Rather than forcing organizations to choose between local and cloud AI, modern AI architectures can intelligently balance both. AI agents can keep routine, high-volume, or data-sensitive workloads local while extending to cloud-based models when additional capabilities are needed.
NVIDIA NeMo Switchyard helps enable this approach as an open source model routing library for AI agents. As agents complete complex workflows involving tasks such as planning, coding, verifying, and summarizing, different steps may benefit from different models. Switchyard automatically routes each workflow step to the most suitable available model, selecting from combinations of open and closed models based on the requirements of the task.
Designed for flexible deployment, Switchyard can be added as a Python library, run as a standalone server, or integrated into an existing agent framework. It works out of the box without requiring additional training or configuration, allowing agents to begin making routing decisions immediately and improve routing performance over time as more workflows are completed.
Switchyard also works alongside NVIDIA NeMo Relay to gain insights into agentic workflows and improve model selection. Relay captures a log of agent activity, including information such as tool calls, model requests, and outcomes, which then allows Switchyard to use these Relay signals to make increasingly effective routing decisions.
By intelligently selecting where workloads are processed, hybrid AI architectures provide organizations with greater flexibility and cost control measures - combining the responsiveness and control of local AI with the expanded capabilities of cloud-based intelligence.
A Flexible Foundation for Enterprise-Ready Agentic AI
With the increasing capability of AI agents, organizations will need infrastructure that can support more flexible approaches to AI deployment. Rather than relying exclusively on cloud-based services or limiting workloads to local resources, organizations, particularly businesses which handle sensitive information or develop innovations, can benefit from an approach that combines both.
Designed for local AI workloads, the Acer Veriton GN100 AI Mini Workstation leverages the NVIDIA Agent Toolkit, including Nemotron 3.5 Lightning and NeMo Switchyard, in delivering leading out-of-the-box accuracy for agentic tasks and can be post-trained for specialized workflows to improve accuracy. It also handily supports a flexible hybrid AI approach - allowing organizations and individuals to run specialized workloads locally while intelligently extending to cloud-based models when additional capabilities are required.
The future of agentic AI won’t be defined by where AI runs, but by how intelligently it can adapt. By combining powerful local compute with flexible model orchestration, organizations and individuals alike can build AI agents that deliver the right balance of performance, control, and capability for every workflow.