New AI PCs from HP and Nvidia aim to slash cloud token costs for enterprises

HP's ZBook Ultra G3a runs large AI models locally, targeting enterprises that want to cut recurring cloud costs. An analyst estimates 20% to 25% of high-end AI workloads will shift to such PCs within two to three years.

Published on: Sep 26, 2026
New AI PCs from HP and Nvidia aim to slash cloud token costs for enterprises

HP has introduced the ZBook Ultra G3a, a mobile workstation designed to run high-end AI workloads locally, a move aimed at enterprises looking to slash recurring cloud token costs. The device, expected in October, uses an AMD Ryzen AI Max Pro processor that tightly integrates the CPU, GPU, and memory on a single chip for faster data movement. A wave of competing laptops powered by Nvidia's RTX Spark superchip is also on the way from Asus, Dell, and Lenovo.

The pitch is straightforward: shift a portion of AI processing from expensive cloud APIs to powerful local hardware. For IT buyers and engineering teams, the math could reshape how they provision tools for AI development, 3D design, and agentic workflows.

Local AI that handles models with billions of parameters

These new machines can run agents, chatbots, video generation, and coding assistants on large models without an internet connection. HP's ZBook Ultra G3a targets enterprise users building, fine-tuning, and running inference on large language models directly on the device. Brian Allen, manager for Global Z workstation products at HP, said the system is built for professionals who are "actually building large language AI models. I'm doing fine tuning. I'm doing inferencing."

Gerardo Delgado, senior director of product management at Nvidia, pointed to the memory and compute headroom as the differentiator. "It's the only PC where you have enough memory to run the large models, enough compute to run them fast, and at the same time, all of your tools that we've accelerated for years… are all working," he said.

Jack Gold, principal analyst at J. Gold Research, estimates that 20% to 25% of high-end AI workloads will run on AI PCs within the next two to three years, rather than entirely in the cloud. The upfront cost of these workstations will be steep-far above Microsoft's Surface laptops that start around $1,200-but Gold said the recurring savings matter most during research and engineering phases that generate multiple project iterations.

Hybrid workflows and the Perplexity partnership

HP has partnered with Perplexity to enable offline AI work in Autodesk Revit, the 3D modeling software widely used in architecture and construction. A designer at an airport without connectivity can run AI locally, then connect to frontier cloud models later to advance the project. Perplexity draws from roughly 19 frontier models to match the task.

Model Context Protocol servers bridge AI models with desktop applications, with user permission. Allen said these connectors will "really start growing and tying into more software applications." The approach keeps sensitive data inside company boundaries while still tapping cloud resources when needed.

Nvidia's Delgado described inference routers as the next major agent capability. These tools redirect work to idle PCs on a network-including Macs and Windows machines with Nvidia GPUs-or to the cloud when a larger model is required. "Everyone that tells you that everything is local is just trying to avoid the fact that the cloud models are improving at an exponential rate," he said.

ROI calculators and buyer caution

HP provides an ROI calculator that lets IT buyers model local versus cloud AI costs. Allen said the tool can show scenarios where "this PC pays for itself in nine months based upon this type of usage." Gold cautioned that such calculators rely on assumptions that may not fit every workplace, though they remain useful for high-level evaluation. In some cases, he noted, absorbing high cloud token costs still makes sense for complex workloads that speed time-to-market.

Deepak Seth, senior director analyst at Gartner, offered a counterweight to the hardware enthusiasm. "You can give them the best tool, and they will still come up with the stupid stuff," he said. The solution, Seth added, is "more power with the right people will lead to better results."

Nvidia saw a surge of customer requests for mobile workstations after models surpassing Opus 4.6 demonstrated strong coding performance. Delgado said development and coding teams are "moving really quickly into local hardware." Agents on these PCs are already driving engineering applications like AutoCAD, though he emphasized that "the idea is not to replace the architect."

Why this matters for IT, development, and construction professionals

For IT leaders and product development teams, the arrival of agentic AI workstations changes the procurement equation. The choice is no longer between thin clients and cloud-only AI. A hybrid model-local processing for iterative work, cloud for frontier model access-can lower token bills and keep proprietary data on-premises. For architects and construction engineers using tools like Revit, offline AI capability means project work continues regardless of connectivity. The key is matching the hardware investment to the skill of the teams using it. As Seth put it, raw compute alone won't deliver results without the right expertise behind it.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)