Archives by Day

September 2026
SuMTuWThFSa
12345
6789101112
13141516171819
20212223242526
27282930

About Rainier

PC gamer, WorthPlaying EIC, globe-trotting couch potato, patriot, '80s headbanger, movie watcher, music lover, foodie and man in black -- squirrel!

Advertising

As an Amazon Associate, we earn commission from qualifying purchases.





NVIDIA Accelerates Local AI With New Agents And NVIDIA PAIR Tool, Reveals Spark PCs Coming In October

by Rainier on Sept. 3, 2026 @ 9:00 a.m. PDT

Local agents get easier to install, faster to run and able to tap multiple RTX PCs at home with NVIDIA PAIR — plus, new NVIDIA RTX Spark Windows PCs arriving in October.

Frontier intelligence is going local. At IFA 2026, NVIDIA, Microsoft and its partners are teaming up to provide faster inference and new tools that make agents easier to set up and run locally on NVIDIA hardware. New compact NVIDIA RTX Spark Windows PCs are also coming in October to give AI enthusiasts, developers and creators more ways to run capable agents locally and securely.

Today's announcements include:

  • Simplified local AI support for NVIDIA GPUs is coming in Hermes Agent, OpenClaw and Perplexity Portable Computer.
  • Up to 1.9x faster local inference — new llama.cpp and vLLM optimizations are available now directly and through LM Studio and Ollama.
  • NVIDIA PAIR — a Personal AI Router tool that intelligently distributes AI inference across the PCs on a user’s local network.
  • NVIDIA RTX Spark arrives in October — with new Windows PCs from Lenovo and Acer. Electronic Arts, Embark and Ubisoft are among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark.

Also, August was a busy month for local AI:

  • Nemotron 3.5 Lightning — which can run on NVIDIA RTX PCs, RTX PRO Workstations, DGX Spark and Jetson — is a 30-billion parameter model that has been launched. Get started with Nemotron 3.5 Lightning today.
  • Z.ai’s GLM-5.3-Flash is a multimodal mixture-of-experts (MoE) model that’s bringing agentic AI to DGX Station.
  • Qwen has released Qwen3.8-Flash-Next, an open weight multimodal MoE model, which can run locally on DGX Spark and DGX Station, along with Qwen3.8-27B, a 27-billion-parameter open model optimized for local agentic and coding workloads on NVIDIA GPUs.
  • LTX’s LTX 2.5 is an open-world video generation model optimized for NVIDIA RTX GPUs, DGX Spark and DGX Station, with new NVFP4, FastVideo and ComfyUI enhancements for faster, more memory-efficient local generation.
  • MiniMax-H3 is an open-weight video generation model with synchronized audio that can run locally on NVIDIA GPUs through ComfyUI.
  • Meta’s Muse Glimmer is a 30-billion-parameter open-weight model for coding and agentic workloads that can run locally on GeForce RTX PCs, DGX Spark, DGX Station and Jetson. NVIDIA has also released NVFP4 quantization with DGX Spark support for more memory-efficient local deployment.
  • DeepSeek v4 Flash is a 284-billion-parameter MoE model with 13 billion active parameters that can run locally on DGX Station.

A Simpler Start for Local Agents

Getting a local agent up and running with local models required some effort — choosing a model, finding a compatible inference server, dialing in quantization settings and keeping everything updated. That friction is disappearing on RTX and DGX systems.

Three of the most widely used agent apps will offer simplified local model setup on Windows, each built on llama.cpp and incorporating NVIDIA’s latest inference optimizations. The new setup experiences are designed to reduce manual configuration and make it easier to get local agents up and running.

Last month, Perplexity introduced its Portable Computer agent, giving users a simple way to run Perplexity locally on Linux systems like NVIDIA DGX Spark with the models, orchestration and tools packaged into a single app experience.

In September, Perplexity Portable Computer will be available on NVIDIA RTX GPUs with at least 24GB VRAM running Linux or Windows, bringing that same streamlined setup to a broader group of PC users. Users can run complete workflows locally without consuming credits, while selectively escalating parts of a task to one of 15+ frontier models in the cloud when additional research or reasoning is needed. Portable Computer asks for permission before sending content to the cloud, helping users keep sensitive information on their device. Here’s some example use-cases:

  • Engineering: Review open PRs in a connected GitHub repo and sort them into ready, blocked, stale, and needs review, each tagged with the next step. Docs that fell out of sync with the latest merge get caught and fixed, with a PR opened for the changes.
  • Finance: Point the agent at two years of brokerage summaries, consolidated 1099s and tax returns and have it trace the recurring holdings creating the most avoidable fees and tax drag, with every figure cited to the exact file and page — all without a document ever reaching a chatbot.
  • Startups: Ask why activation went flat, and the agent analyzes the funnel export locally to find where new signups drop off between install and first completed task, then posts the top insights straight to the team's Slack channel.

Hermes Agent — developed by Nous Research — is a general-purpose agent used by millions that excels at reliability and self-improvement. Model- and provider-agnostic, Hermes is built to run all day on local systems, making RTX PCs, RTX PRO workstations and DGX Spark a natural fit.

Coming soon, configuring a local model in Hermes will provide users with one-click setup across RTX and DGX systems on both Windows and Linux. The agent will automatically detect the NVIDIA GPU, select an appropriate model and configuration, and run it through integrated llama.cpp with NVIDIA inference optimizations already in place, eliminating manual model downloads and tuning.

Once it is running, Hermes works the way it does anywhere else. It uses tools, maintains context across tasks, remembers information between sessions and creates reusable skills over time, allowing the agent to become more capable with continued use. Running the model locally on a GPU keeps performance fast while keeping data on the system.

One-click local model setup is coming soon. Learn more about Hermes Agent.

OpenClaw has become one of the defining projects of the open-agent movement — the largest AI project on GitHub, with more than 380K stars and a fast-growing community that’s building tools and skills across research, engineering, project management and everyday productivity.

NVIDIA, Microsoft and OpenClaw have been working together to make that experience easier to set up on Windows PCs. To reduce onboarding friction, the OpenClaw Windows App simplifies the process of setting up an optimized local model on any RTX GPU with at least 24GB of VRAM.

Faster Inference Gives Local Agents a Boost

Inference performance is critical to keeping local agents responsive. NVIDIA is continuing to collaborate with the open-source llama.cpp and vLLM communities to accelerate agentic workloads across local NVIDIA platforms.

llama.cpp delivers up to 1.9x higher throughput through kernel optimizations on a GeForce RTX 5090, enhanced speculative decoding techniques and faster prefill.

vLLM delivers 1.2x on RTX PRO 6000 Blackwell Workstation Edition and up to 1.4x on two DGX Spark clusters. New XQA attention kernels in FlashInfer and backend optimizations help to accelerate inference across both platforms.

Experience these gains now through LM Studio, Ollama, llama.cpp and vLLM.

Tap Idle PCs for More Local AI Compute With NVIDIA PAIR

More than half of U.S. households have two or more PCs, and much of that computing power sits idle throughout the day. NVIDIA Personal AI Router (PAIR) is a free, open source software tool that puts those systems to work together for local AI.

Agentic workflows often break complex tasks into smaller jobs that can run in parallel, but performance can slow when every request is competing for the same GPU. PAIR automatically discovers compatible PCs on a local network and routes independent inference requests to whichever system has capacity. It works with Ollama and LM Studio and can adapt as devices join or leave the network.

For example, a user could ask Hermes to create a “Sunday Reset” plan by sorting through a cluttered inbox and prioritizing what needs attention now, what can wait and what can be skipped. Hermes can split that work across multiple subagents, while PAIR distributes those jobs across available PCs instead of having them all wait on a single GPU.

The result is more compute for local agents, with more tasks running in parallel and the flexibility to move AI workloads to another PC while the main system is being used for gaming, creating or other work.

The NVIDIA PAIR beta is available for Windows, macOS and Linux through both graphical and terminal interfaces, supporting NVIDIA GeForce RTX 20 Series GPUs and newer, NVIDIA RTX PRO workstation GPUs (Turing architecture and newer), NVIDIA DGX Spark and Apple M4 or newer silicon.

Powerful On Device Photo Editing With Cyberlink PhotoDirector AI PC Mode on RTX Spark

Open image and video models enable artists to experiment with Creative AI models on PCs. This enables artists to iterate and explore concepts and ideas, without the dreaded token anxiety and keep more of their creative work private and on-device.

CyberLink’s new PhotoDirector AI PC Mode is one of the first applications to integrate these diffusion models directly into a creative software, and turn them into a creative tool at the finger tips of the artists. Coming to PhotoDirector 365 and optimized for NVIDIA RTX Spark when it launches, AI PC Mode users are getting AI-powered editing tools for generative editing, image enhancement, object and distraction removal, background removal and replacement, portrait refinement and the creation of entirely new visuals — with the flexibility to choose between local or cloud processing, depending on the task.

On NVIDIA GPUs, PhotoDirector uses TensorRT-RTX and FP8 to accelerate local AI.

Start using Cyberlink’s PhotoDirector 365 photo editing software and learn more about PhotoDirector AI PC Mode, launching with RTX Spark in October.

NVIDIA RTX Spark Windows PCs Arrive October 2026

NVIDIA RTX Spark is coming this October— and at IFA 2026, partners are showing off their hardware. At IFA, newly announced designs join the existing six OEMs shipping in October. Acer showed its compact desktop RTX Spark concept, and Lenovo announced its Yoga Pro 9n and Yoga 9n 2-in-1.

RTX Spark is a new beginning for Windows PCs. One PC built for creators, gamers and AI agents. With a powerful 1 Petaflop RTX Blackwell GPU, up to 128GB of unified memory and a highly efficient 20-core Grace CPU, RTX Spark delivers incredible performance and efficiency. This superchip enables high performance thin laptops with all day battery life and compact desktops to power always-on agents. Paired with the new Windows Agent framework, it enables agents that run safely in the background under OS level control.

Last week at Gamescom, Electronic Arts, Embark and Ubisoft were among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark Windows PCs. They join the publishers that announced RTX Spark support at COMPUTEX in May, including KRAFTON, NetEase, Riot Games and XBOX.

Sign up to be notified when RTX Spark laptops and desktops are available.

ICYMI: More Updates From NVIDIA Local AI

  • NVIDIA Brings New RTX Tech and Games to Gamescom — NVIDIA released DLSS 4.5 Ray Reconstruction, featuring a new second-generation transformer model for improved image quality in ray-traced and path-traced games. Gamescom also brought new RTX announcements for titles including 007 First Light, CONTROL Resonant and Gears of War: E-Day, plus expanded game support for the upcoming NVIDIA RTX Spark.
  • Introducing DeepSeek Harness — DeepSeek’s new open source harness pairs with DeepSeek-V4-Flash to power local agentic coding workflows on NVIDIA DGX Station and multi-DGX Spark setups.
  • MLPerf Client v2.0 Expands AI PC Benchmarking — MLCommons released MLPerf Client v2.0, developed in collaboration with NVIDIA and other industry leaders. The update adds new benchmarks for agentic AI and image generation, alongside expanded LLM testing for real-world local AI workloads.

Put Your Local AI Systems to Work with NVIDIA PAIR

AI agents are learning to do more by working together. A lead agent can break a complex task into smaller jobs and assign those jobs to specialized subagents. This breadth-first approach can improve the speed of task completion, but it also can bottleneck the system as many requests are sent to the GPU simultaneously.

NVIDIA Personal AI Router (PAIR) leverages your local hardware to relieve this subagent bottleneck. PAIR routes each independent inference request to an available system on the home network. It works with familiar local inference services, including Ollama and LM Studio, so users can expand the compute available to an agent without redesigning the agent itself. No agent harness changes are necessary.

The NVIDIA PAIR beta is available for supported Windows, macOS, and Linux systems through graphical and terminal interfaces. It supports compatible systems with NVIDIA GeForce RTX 20 Series GPUs and newer, as well as NVIDIA DGX Spark and Apple M4+ silicon.

What is NVIDIA PAIR?

The product soul of NVIDIA PAIR is simple: maximize the AI compute in your home.

PAIR is a virtual inference router, not a new inference engine. Ollama or LM Studio still runs the model on a selected machine. PAIR discovers participating systems, tracks whether each system is ready for a request, schedules independent jobs and returns each response to the application that originated it.

Agents can send a request through the familiar local interface it expects. PAIR receives the request through its proxy, identifies its engine and model requirements, and selects one eligible node. That node executes the request from start to finish and sends the response back through PAIR. The agent continues to see one connection while PAIR handles placement behind it.

  • No new API: PAIR proxies compatible Ollama and LM Studio interfaces rather than asking every agent harness to integrate with a new cluster API.
  • Elastic clients: Compatible systems can contribute capacity when available and drop away when needed such as powering down or hibernating the system.
  • Local control: PAIR is designed to keep prompts, data and inference traffic on the user's existing local network.

How are subagents changing the local inference problem?

Consider an AI prosumer running a local agent on a primary RTX AI PC. The agent receives a research, coding or personal-organization task and divides it among several sub-agents. Each worker explores a bounded part of the problem while other workers verify evidence or assemble the result.

From the user's perspective, this is one task. At the inference layer, it can become dozens of independent model calls. If every call targets one local engine, they compete for the same execution slots. The queue grows, and the primary PC stays occupied even though an RTX workstation, laptop or DGX Spark elsewhere on the network may have compatible capacity available.

PAIR lets the inference layer go wider with the agent. Some sub-agent requests can run on the primary PC while others run on additional paired nodes. When the workload has enough independent work, using more ready systems can reduce queueing and improve end-to-end completion time.

It can keep the primary PC focused on graphics-intensive gaming, content creation or other interactive work while distributing inference to other nodes.

This is workload-level concurrency. PAIR does not make one inference request run across several GPUs. Every request is assigned to one eligible node and remains there for its lifetime.

Home AI clusters are elastic by design

A home cluster is not a miniature data center. Dedicated clusters are typically built around systems that remain powered, consistently configured and continuously available. Home hardware is dynamic. A gaming PC may become busy playing a game. A laptop may sleep, close, or leave the network. A workstation may have the requested model while another machine does not. An inference engine can be stopped, or the user can reclaim a GPU for a foreground application.

PAIR is designed around those changing conditions. It can discover local systems with mDNS, pair supported devices on the private network and maintain a live view of which nodes can accept new work. Client nodes can join the available pool when ready and drop away when needed without turning the home into a dedicated, always-on inference installation. For each new request, PAIR considers factors including:

  • Whether a paired node is online and ready
  • Whether a supported inference engine is enabled
  • Whether the exact requested model is present
  • The current node and engine workload, including active jobs
  • The existing GPU utilization (ie, if there is a graphics intensive app or tool running)

PAIR does not depend on every system being identical or permanently available. It schedules inference requests and manages the cluster around how the systems are used in every-day life.

Demo: Hermes subagents go wider

This demonstration pairs PAIR with Hermes Desktop, which creates the sub-agent workload, and Ollama, which executes the model on each selected node. The task asks Hermes to analyze a synthetic household inbox and produce a trustworthy Sunday Reset plan: what must happen tonight, this week, later or not at all, with evidence for each material conclusion.

For the five-subagent run, Hermes creates 5 specialists to review independent parts of the evidence, reconcile conflicts and return a consolidated plan. Hermes owns decomposition, delegation and synthesis. PAIR owns inference routing. Ollama executes each request on the node PAIR selects.

Using Qwen3.6 35B A3B on one RTX 5090, the same five-subagent workload took 6 minutes 18 and seconds to complete on average, while a two-device PAIR cluster containing two RTX 5090s took 3 minutes and 48 seconds to complete on average. This is an unofficial, configuration-specific demonstration - not a general benchmark or a promise of linear scaling.

Hermes three-subagent demo results

The Jobs view in PAIR is the ground truth for where inference ran. Hermes agent count and PAIR job count are different measurements because one agent can generate multiple model requests. Multi-node execution should be claimed only when PAIR telemetry shows jobs running on more than one eligible node.

How does NVIDIA PAIR work?

1. Install, discover and pair systems

Install PAIR on each compatible Windows, macOS, or Linux system. PAIR uses local-network discovery (mDNS) to find nearby systems automatically; a node can also be added by IP address when needed. The user approves a secure pairing request to create the trusted set of local nodes PAIR can consider.

All node to node communication is blocked until the secure connection and pairing is established. Once the connection is made, communications are secured with MTLS and generated certificates so that the communications between the nodes stay private on the network.

2. Prepare inference engines and models

Each participating node runs a supported local inference engine: Ollama or LM Studio. PAIR can help install an engine and initiate model downloads on paired systems, reducing the work required to prepare several machines. A node becomes eligible for a request only when the required engine is enabled and the exact requested model is available there.

However, models do not have to be identical across each node in cluster. Different systems can host different models, and PAIR can route according to model location. Loading the same model tag on more nodes simply gives the scheduler a larger eligible pool of nodes for that request.

3. Proxy the compatible local interface

A compatible application sends an Ollama-compatible or LM Studio-compatible request through the local endpoint proxied by PAIR. Agent harnesses can continue using the interface they already understand instead of discovering and integrating with every machine independently. Applications that expose a configurable base URL can continue to point at their Ollama or LM Studio respective endpoints.

PAIR proxies the request by taking over the default port that Ollama and LM Studio use for their services. If the agent harness is using a different port, the proxy port can be configured in the PAIR engine settings.

PAIR inspects the request's engine and model requirements, then passes those requirements to the router. This separation is central to the design: the agent decides what work to request, while PAIR decides where eligible work should run.

4. Schedule one eligible node

The scheduler filters paired systems using current information about readiness, supported engine state, requested-model presence and job load. It selects one eligible node, and the PAIR router on that system passes the request to the local inference engine. Independent calls from other sub-agents can be assigned to other ready nodes at the same time.

5. Return the response and make routing visible

The selected engine executes the request, and PAIR streams the response back through the same local interface to the originating application. The Jobs and metrics views show which node handled each routed request, making placement observable.

What PAIR Does - and Does Not Do

PAIR is most useful for workloads that expose several independent requests at the same time, including multi-agent applications and concurrent local AI tools.

For those workloads, PAIR can:

  • Route independent jobs across ready systems on the local network.
  • Reduce queueing when several requests would otherwise wait behind one local engine.
  • Improve completion time for a suitably parallel workload in a compatible configuration.
  • Help free the primary PC for gaming, creation or other interactive tasks.
  • Keep the application workflow familiar and local-first.

PAIR does not:

  • Merge GPUs or pool VRAM into one larger accelerator.
  • Shard a single model or split one inference request across machines.

Highly sequential tasks, workloads dominated by one long model call or configurations in which only one node has the requested model may see less benefit. Measure the workload that matters using end-to-end completion time, queueing, output quality and observed routing on the actual systems in use.

Get started with NVIDIA PAIR

PAIR brings a cluster-like experience to the dynamic NVIDIA systems already in the home while keeping the local inference workflow familiar.

To get started:

1. Download the NVIDIA PAIR beta for a supported Windows, macOS, or Linux system.
2. Install PAIR on the NVIDIA RTX PCs, workstations or DGX Spark systems to include.
3. Discover and securely pair the systems on the local network.
4. Enable Ollama or LM Studio and download/place the required models on eligible nodes.
5. Run a compatible agent on a system with PAIR installed that uses Ollama or LM Studio.

The NVIDIA PAIR project is open source, and developers can inspect the code, report issues and contribute improvements to discovery, pairing, routing, engine integration, models, endpoints and the user experience.

Install. Pair. Run - and put your home AI systems to work.

Tune into NVIDIA Local AI to explore what’s possible with local AI and discover the latest open models, development frameworks and agentic tools running across NVIDIA hardware, including RTX PCs, RTX PRO workstations, DGX Spark and DGX Station.

blog comments powered by Disqus