Perplexity Put an AI Agent on Your GPU. The Cloud Gets Visitation Rights.

Perplexity’s Portable Computer runs AI agents locally for privacy and lower costs—provided your “personal computer” has at least 24GB of VRAM.

Share
SiliconSnark robot guards a local AI computer while a cloud asks permission to enter.

The most persuasive AI demo of the week may be a billing counter that does absolutely nothing.

In Perplexity’s demonstration of its new Portable Computer, an agent chewed through a folder of tax forms and investment documents while the little cloud-credit meter sat at zero, experiencing the sort of spiritual peace normally reserved for unplugged parking meters. The files stayed on the machine. The model stayed on the machine. The work stayed on the machine. Somewhere, a hyperscaler felt a brief chill and blamed the air conditioning.

Announced August 25 in partnership with Nvidia, Portable Computer takes Perplexity’s agentic Computer platform and packages the model, inference engine, agent harness, tools, connectors, and security sandbox into a local app. It launches on Linux for Perplexity Pro, Max, Enterprise Pro, and Enterprise Max subscribers. You need an Nvidia GPU with at least 24GB of VRAM—roughly an RTX 3090 or newer—or a DGX Spark. Windows support is due in September. Mac support, despite the local-AI community’s ongoing religious relationship with Mac minis, is not on the roadmap.

This is not a cloud chatbot with a privacy toggle tastefully buried under six menus. Every task begins locally. If the smaller on-device model runs into a problem requiring heavier reasoning, it can ask permission to send a limited step to a frontier model in the cloud. The remote model returns advice but never receives direct access to local files or tools.

That is a genuinely interesting product decision. It is also an excellent summary of modern adulthood: do as much as you can at home, then call an expensive specialist and reveal only the embarrassing part.

Your PC Is Local, Assuming It Resembles a Minor Research Lab

“Portable” is doing some reputational heavy lifting here. Portable Computer does not mean you can install it on the laptop currently wheezing through 27 browser tabs and a video call. The 24GB VRAM floor excludes most consumer PCs. Nvidia’s DGX Spark, meanwhile, is a compact desktop AI box, but “compact” in this context means the supercomputer has agreed not to occupy an entire room.

The initial audience is therefore developers, AI enthusiasts, and privacy-sensitive companies that already own serious Nvidia hardware. This is not yet the democratic return of personal computing. It is more like personal computing with a velvet rope and a CUDA installation.

Still, the hardware caveat should not obscure the shift. We have spent years watching the AI industry centralize intelligence inside enormous remote clusters. The standard bargain was simple: send your data to somebody else’s warehouse, pay by the token, and trust a privacy page written with the emotional warmth of a customs form. Portable Computer reverses that default. Your machine does the routine work; the cloud becomes an exception that needs permission.

That matters most for the kinds of jobs agents are actually being asked to do: reviewing contracts, reconciling financial documents, analyzing source code, sorting internal research, and wandering through folders whose filenames tell the full story of an organization’s governance. Our guide to computer-use agents argued that convenience and capture are usually roommates. Running the agent locally does not evict risk, but it at least gives capture a lease agreement.

The Harness Is the Product, Not Decorative Agent Confetti

The clever part is not merely downloading an open model. People have been doing that for years, usually after a tutorial promised “five easy minutes” and then introduced Docker, quantization, drivers, an inference server, three incompatible repositories, and a forum reply from 2024 that ends with “fixed it.”

Portable Computer packages the full operational stack. Perplexity calls the surrounding machinery the agent harness: the prompts, tool definitions, context management, execution loop, verification steps, permissions, and sandbox that turn a model from a talker into something capable of completing multi-step work.

In the technical paper released with the product, Perplexity makes a smart argument: smaller local models cannot simply be dropped into harnesses designed for giant frontier systems and expected to thrive. The company keeps its core prompt and toolset small, loads specialized skills only when needed, compacts stale context, and converts bulky connectors into lighter command-line tools. It also disables tool use if the OS-level sandbox is unavailable rather than quietly running commands with the user’s full permissions. More agent products should possess this basic instinct for not setting the curtains on fire.

Perplexity offers Qwen 3.8 27B and its post-trained PPLX 27B at launch, with Nvidia’s Nemotron 3.5 Lightning coming later. Its own Local Knowledge Work Bench says the base Qwen model with Perplexity’s harness scored 82.6%, ahead of Pi at 77.6% and Hermes at 74.0%; PPLX 27B reached 85.4%. Those are company-produced results on a 53-task benchmark the company says it plans to publish, so the usual benchmark caution lights remain illuminated. A vendor winning a vendor-designed test is evidence, not canonization.

But the underlying principle is sound. We saw something similar when SAP built agents around the systems where business work already lives. Models matter. The boring connective tissue that constrains them, feeds them context, and checks their work often matters more.

Free Tokens, Plus the Electricity You Have Agreed Not to Mention

Perplexity and Nvidia describe local inference as having essentially zero marginal token cost. That is true in the same way a home-cooked meal has no restaurant bill. You still bought the kitchen.

The GPU costs money. Electricity costs money. Setup and maintenance cost time. A subscription remains involved. “Zero token costs” means no per-token cloud meter for work completed locally, not that computation has escaped thermodynamics and joined a cooperative.

Yet this distinction becomes important for agents because agents are token furnaces. A chatbot answers and stops. An agent reads, plans, calls a tool, examines the result, revises the plan, verifies the output, and occasionally spends 40,000 tokens rediscovering that a file was in the Downloads folder. If we are serious about always-on software workers, paying retail cloud rates for every internal thought becomes an ugly operating model.

This is why the launch is more consequential than another local-chat app. In our investigation into whether AI agents make money or merely create Mac mini lifestyle photography, the recurring answer was that useful agents need permissions, memory, logs, limits, and economics that survive contact with reality. Portable Computer directly attacks two of those realities: recurring inference cost and the governance problem of sending every private token off-device.

The Cloud Still Gets a Key, but It Has to Knock

Compact models remain weaker on hard reasoning, and Perplexity does not pretend otherwise. On Terminal Bench 2.1, its local Qwen setup scored 59.6%. Allowing the local agent to consult Claude Opus 5 raised that to 73.0% at an estimated $0.415 per task. Claude alone scored 82.4% at $0.65.

The hybrid system therefore recovers a meaningful chunk of frontier performance while preserving local control over tools and files. Before an advisor call, the harness selects relevant context, runs a personally identifiable information classifier, and shows the user what would leave the device. The user can reject it. The advisor sees only the approved context and returns text guidance.

This may be the product’s most important idea. “Local AI” is too often marketed as a purity test: either every electron stays under your desk or you have surrendered to the cloud empire. Real work is messier. Sometimes a 27-billion-parameter model is enough to extract tables from PDFs. Sometimes you need the expensive oracle. A system that understands the difference—and asks before crossing the boundary—is more useful than either ideological extreme.

It also gives Nvidia a tidy strategic win. The company sells the hardware that powers the cloud, the hardware under your desk, and much of the software connecting them. Nvidia is not choosing between centralized and local AI. It is selling shovels to both sides of the property dispute.

The Verdict: A Real Shift Wearing Enthusiast Hardware

Portable Computer is not the AI PC for everyone. It is Linux-first, Nvidia-only, subscription-gated, benchmarked largely by its maker, and dependent on hardware most people do not own. The cloud fallback also means privacy rests on understandable approvals, careful context selection, and a PII classifier that will need to perform reliably when documents get strange.

But this is not an expensive vibes machine. It is a meaningful incremental move toward agents with sane economics and explicit data boundaries. The product treats local models as capable workers with known limitations, frontier models as consultants rather than landlords, and the harness as engineering rather than garnish.

I am impressed, which is inconvenient. Perplexity has taken the local-agent fantasy—your data, your machine, your always-on robot employee—and added the part most fantasies omit: a permission dialog, a sandbox, and an honest admission that sometimes the small model needs to call an adult.

The future of personal AI may indeed live on your computer. For now, your computer just needs 24GB of VRAM, Linux, a paid plan, and the electrical temperament of a modest space heater. Progress rarely arrives without accessories.