Perplexity’s new AI stack runs on NVIDIA’s DGX Spark computer

Perplexity is now running its entire system on NVIDIA’s DGX Spark computer, integrating the agent harness, orchestrator, and post-trained models onto this hardware. The system allows agents to run locally while accessing over 15 cloud models, but only after the orchestrator requests permission before routing a task to them, prioritizing user control, Portable Computer says. Portable Computer currently utilizes PPLX 27B or Qwen 27B for on-device AI processing, with Nemotron 3.5 Lightning scheduled to be added soon, expanding the capabilities of local AI workflows.

Local Agent Stack Deploys on NVIDIA DGX Spark

Portable Computer is now deploying its complete local-first agent stack on NVIDIA DGX Spark systems, a move indicating confidence in the hardware’s capacity for advanced artificial intelligence workflows. This architecture prioritizes local processing for speed and privacy, engaging cloud models only after explicit approval. Each task initiated within Portable Computer executes locally on the DGX Spark; however, the orchestrator functions as a gatekeeper, ensuring user consent before routing any task requiring extensive reasoning to one of over 15 integrated cloud models.

This approach balances the benefits of on-device AI with the power of remote computation, giving users control over data handling and processing location. A demonstration of this process involved reviewing pull requests on GitHub, where the local orchestrator grouped, tagged, and prepared a summary for posting to a Slack channel, all while maintaining data within the DGX Spark environment.

Portable Computer currently utilizes Qwen 3.8 27B and PPLX 27B for on-device AI processing, with Nemotron 3.5 Lightning coming soon, according to the company. The system’s sandbox infrastructure isolates code and tool execution, providing controlled access to local files and connected applications like Gmail, Outlook, and Slack, the company says. Work completed by the local models incurs no per-token charge, making large-scale operations such as repository migrations and batch file summarization economically viable on owned hardware.

The installation process sets up the local agent harness, orchestrator, and post-trained models with sandboxed execution. Available to Pro and Max subscribers, the system requires a DGX Spark with GB10, 128 GB of memory, and at least 1 TB of storage to function optimally.

Stay current

See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.

Avatar of Rusty Flint

Rusty Flint

Rusty is a quantum science nerd. He's been into academic science all his life, but spent his formative years doing less academic things. Now he turns his attention to write about his passion, the quantum realm. He loves all things Quantum Physics especially. Rusty likes the more esoteric side of Quantum Computing and the Quantum world. Everything from Quantum Entanglement to Quantum Physics. Rusty thinks that we are in the 1950s quantum equivalent of the classical computing world. While other quantum journalists focus on IBM's latest chip or which startup just raised $50 million, Rusty's over here writing 3,000-word deep dives on whether quantum entanglement might explain why you sometimes think about someone right before they text you. (Spoiler: it doesn't, but the exploration is fascinating)

Latest Posts by Rusty Flint: