Nebius has acquired Inferize, an inference optimization startup, to cut the time and cost required to launch and scale large AI models on its Nebius Token Factory platform. The deal, announced October 1, 2026, brings technology that directly addresses the "idle GPU tax" - the wasted compute and expense when models sit cold before serving requests, during demand spikes, or when weights are updated mid-run.
Inferize's team and technology will fold into Nebius Token Factory, the company's managed inference platform for production AI. The acquisition adds another layer to Nebius's inference stack, which already includes model-level and kernel-level optimizations from Eigen AI and system-level orchestration from Clarifai's core team and licensed technology.
The cold start problem
Running inference at scale creates a persistent headache for operations teams. Cold starts - the time models need to load before they can serve a single request - leave assigned GPUs idle. This happens at launch, when demand spikes force new instances to spin up, and during mid-run weight updates in processes like reinforcement learning. Platforms often hold spare capacity just to meet service-level targets, burning money on hardware that sits waiting.
Inferize's technology reduces this idle time. Capacity can scale more closely with actual usage, driving higher utilization and better token economics. That translates to more customer demand served from the same GPU footprint.
What leadership said
Danila Shtan, Chief Technology Officer of Nebius, said: "Running inference well takes more than fast GPUs and optimized models. The whole system needs to respond when demand changes, including how quickly additional capacity is ready to serve customers. Inferize brings technology that accelerates that process and a team with deep expertise in GPU systems."
Guy Bortnikov, co-founder and CEO of Inferize, said: "Keeping spare GPUs running is the price of being ready for demand. Removing that cost is what we built Inferize to do, and Nebius is where it can go straight into the platform."
Team and timeline
Inferize was founded in January 2026 and had a working prototype within three months. The engineers will work across Nebius Token Factory, starting with the integration of their technology. The terms of the transaction were not disclosed.
For executives and technical leaders building internal AI capabilities, the core challenge Inferize addresses is a familiar one: inference infrastructure that sits idle is a direct drain on margin. Courses on AI Technology Leadership cover the infrastructure decisions that shape unit economics like these.
Why this matters for operations and strategy leaders
Every percentage point of GPU utilization gained or lost flows straight to the bottom line in production AI. Inferize's approach targets the gap between provisioned capacity and actual throughput - a gap that grows wider as models get larger and demand patterns become less predictable. For teams running inference at scale, technology that shrinks cold start latency reduces the buffer capacity needed to meet SLAs. That means fewer idle GPUs, tighter token costs, and infrastructure that flexes with demand rather than over-provisioning against it.
Your membership also unlocks: