At IFA 2026, NVIDIA and Microsoft, along with partners, are pushing local AI onto the desktop with a new line of compact Windows PCs called RTX Spark, arriving in October from Lenovo and Acer. Game publishers Electronic Arts, Embark, and Ubisoft are bringing titles to the platform. The machines are designed to run AI models locally, and NVIDIA says its llama.cpp optimizations deliver as much as 1.9 times higher throughput using a GeForce RTX 5090, while vLLM shows smaller gains on workstation and DGX Spark setups, with both reaching users now in LM Studio and Ollama.
NVIDIA also introduced PAIR, the Personal AI Router, a free, open-source beta that distributes inference jobs across PCs on a local network. PAIR works with Ollama and LM Studio and supports GeForce RTX 20 Series and newer, RTX PRO workstation GPUs from Turing onward, DGX Spark, and Apple M4 or newer. The software automatically finds compatible machines and routes requests to whichever has capacity, adapting as devices join or leave. The company notes that more than half of U.S. households have two or more PCs, making such distributed inference practical.
Read more: Nvidia Announces RTX Spark Processor Family for Laptops and Desktops at Computex 2026
Several new open-weight models were announced for local use. Nemotron 3.5 Lightning is a 30-billion-parameter model that runs on RTX PCs, RTX PRO, DGX Spark, and Jetson. Meta's Muse Glimmer, also 30 billion parameters, targets coding and agentic tasks and runs on GeForce RTX PCs, DGX Spark, DGX Station, and Jetson. DeepSeek v4 Flash, a 284-billion-parameter mixture-of-experts model with 13 billion active parameters, runs on a two-node DGX Spark cluster or DGX Station. Qwen released Qwen3.8-Flash-Next, an open-weight multimodal model that runs locally on DGX Spark and DGX Station, plus a 27-billion-parameter Qwen3.8-27B for local agentic and coding work.
For video generation, LTX 2.5 is optimized for RTX GPUs, DGX Spark, and DGX Station, while MiniMax-H3 generates video with synchronized audio, and its FastH3 distilled version claims a sevenfold performance improvement. Z.ai's GLM-5.3-Flash multimodal model brings agentic AI to DGX Station.
Local AI assistants also got simpler setup. Hermes Agent, developed by Nous Research and used by millions, now offers one-click local model setup on Windows, with Linux support coming soon. OpenClaw, described as the largest AI project on GitHub with more than 380,000 stars, has a Windows app that simplifies local model setup on RTX GPUs with at least 24GB of VRAM. NVIDIA, Microsoft, and OpenClaw collaborated on the Windows setup. Perplexity's Portable Computer is available on RTX GPUs with at least 24GB of VRAM on Linux, with Windows support coming soon, and can escalate to more than 15 frontier models in the cloud.
The announcements build on NVIDIA's recent push into AI infrastructure, including a deepened partnership with MediaTek for edge-to-cloud platforms and a collaboration with CrowdStrike on agentic cybersecurity.













