Complete AI Training

AI news ·

Prime Intellect launches Prime Inference for serving frontier open models

Prime Inference processes nearly a trillion tokens daily and uses separate GPU groups for prefill and decode, cutting p90 inter-token latency by nearly 40%.

Share

Prime Intellect released Prime Inference on October 2, 2026, a serving platform for open-source AI models that already processes nearly a trillion tokens daily across internal workloads and large-scale customer deployments. The release closes a gap in the company's continuous learning loop, connecting trained models to real users and feeding production data back into training.

The platform offers serverless endpoints and reserved capacity on NVIDIA Blackwell GPUs across multiple data centers, with Vera Rubin hardware coming soon. Its first public deployment, GLM-5.3, went live on OpenRouter on September 22 and has maintained 100% uptime with a near-zero tool-call error rate.

Serving infrastructure built for reliability

Prime Inference separates its public API from the model fleet so capacity can move, fail, or scale without changing client endpoints. Shared circuit breakers let gateway replicas react to failures consistently, while lease-based admission control prevents overload and recovers capacity when a process disappears. Every cluster undergoes continuous monitoring with health checks reaching down to NVLink and InfiniBand.

Alerts go to an on-call team staffed 24/7. The company maintains significant overflow capacity to reroute traffic and bring up new deployments when needed. "Our stack combines NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer, developed in close partnership with Inferact and NVIDIA, with improvements contributed upstream," the team said.

Disaggregated prefill and decode

Prime Inference runs prefill and decode on separate GPU groups to prevent long prompts from interrupting token generation for existing sessions. NVIDIA Dynamo handles routing and orchestration, while vLLM runs the model on each group. Once prefill finishes, the decoder pulls computed KV through NIXL and adds the request to its batch. This approach reduced p90 inter-token latency by nearly 40% in the company's tests.

Dynamo's KV-aware router selects a prefill worker based on cached prompt availability and queued work. Mooncake provides a second cache tier in host DRAM, allowing prefixes offloaded from GPU memory to be retrieved instead of recomputed. This lets the system retain more conversation history across sessions.

Performance tuning for agent workloads

A typical agent turn adds about 6,000 tokens to a 140,000-token prompt, reusing most conversation history. Prime Inference tuned GLM-5.3 on GB200 NVL72 for interactive speed, model quality, and concurrency. The interactivity target was 100 end-to-end tokens per second per user, and the team optimized prefill topology, scheduling, KV compression, and transfer paths to hit that bar.

NVFP4 compression reduced each MLA cache row from 576 to 352 bytes, increasing total cache capacity by roughly 50%. A native sparse-MLA kernel built by the team eliminates intermediate GPU memory buffers, unpacking compressed rows on-chip as attention needs them. At 15 query tokens per launch, the optimized kernel took approximately 12.0 microseconds on GB200, compared with 17.7 microseconds for the staged path.

For KV transfer, the team adopted the BLHNC block-major layout from the vLLM community. This reduced transfer descriptor count by roughly 10× and lowered mean transfer time by about 47%, from 146 milliseconds to 78 milliseconds, compared to the layer-major layout.

Why this matters for operations and technology leaders

Prime Inference demonstrates that disaggregated serving with automatic failover can deliver production-grade reliability for open-source models at scale. For executives evaluating AI infrastructure, the architecture's separation of public API from model fleet means capacity decisions and hardware failures do not disrupt client applications. The platform's unified billing and team-level usage tracking also address a common pain point: managing inference spend across multiple models and workloads. Technology leaders building agent systems should note the tool-call reliability work - Prime Intellect contributed a structural-tag builder to Dynamo and fixed parsing errors that caused silent failures, resulting in near-zero tool-call error rates during long-running agent sessions.

Share