Gimlet Labs and Cerebras Systems plan to bring ultrafast AI inference to the cloud, targeting speeds up to 3,000 tokens per second for demanding agentic and real-time applications. The first Cerebras-powered Gimlet Cloud datacenter is expected to come online later this year, combining Cerebras' wafer-scale compute with Gimlet's disaggregated inference architecture across infrastructure and developer APIs.
The speed of token generation shapes what developers can build and how users experience AI. For real-time voice, video, and agent-based applications, latency determines whether an interaction feels fluid or sluggish. When AI responds in real time, users engage more, stay longer, and run higher-value workloads. Fast tokens are more valuable tokens.
How the integrated inference stack works
Gimlet Cloud combines the Cerebras Wafer Scale Engine with GPUs into a single inference solution. The platform uses advanced inference disaggregation technology to orchestrate model execution so each phase runs on the silicon best suited to it. The result is a purpose-built cloud optimized for production-scale agentic and real-time workloads.
"Inference speed matters. It determines how productive AI can be. Fast inference creates magical user experiences and opens new markets," said Zain Asgar, co-founder and CEO of Gimlet Labs. "By combining Gimlet's multi-silicon software with the Cerebras Wafer Scale Engine, we can run each phase of inference on the hardware best suited to it and plan to deliver up to 3,000 tokens per second at production scale."
Sean Lie, co-founder and CTO at Cerebras, framed the partnership in terms of datacenter economics. "Combining the fastest tokens from Cerebras with the highest throughput GPUs delivers the best datacenter economics for everyone," Lie said. "Everyone wants more high value tokens. Cerebras delivers the fastest AI inference in the world, and GPUs deliver high throughput. By making Cerebras a native part of its inference cloud, Gimlet will bring our industry leading speed and intelligent AI to more developers at production scale."
From private deployments to public cloud
The collaboration builds on joint customer engagements underway since last year, with an integrated solution already serving tokens in private deployments. Gimlet Labs will now expand the work to integrate software, infrastructure design, APIs, developer tooling, optimization, validation, and production operations to make ultrafast inference broadly available through Gimlet Cloud.
Gimlet Labs is backed by Andreessen Horowitz and Menlo Ventures and focuses on breakthrough improvements in AI performance through techniques such as automated GPU kernel generation, workload orchestration, and heterogeneous execution across diverse hardware. Cerebras, which trades on NASDAQ under the ticker CBRS, builds wafer-scale AI infrastructure used by global corporations, research institutes, and governments.
Why this matters for IT, operations, and research teams
For teams running production AI workloads, inference speed directly affects infrastructure costs and user retention. A system delivering 3,000 tokens per second changes the calculus for real-time applications that cannot tolerate lag - voice assistants, video processing pipelines, and autonomous agents. The disaggregated approach, where different phases of inference run on different silicon, also means teams can optimize hardware allocation rather than over-provisioning on a single chip type. Developers and technical leaders evaluating inference architectures can explore AI for Developers Courses to understand how heterogeneous compute strategies affect production systems. For CTOs and engineering leads planning multi-year infrastructure roadmaps, the shift toward wafer-scale inference disaggregation represents a concrete architectural decision point worth tracking as the first Cerebras-powered Gimlet Cloud datacenter comes online.
Your membership also unlocks: