AI news ·
Ai2 overhauls GPU cluster scheduling with budget-based system to reduce squatting and on-call toil
Ai2 replaced its GPU priority scheduler with a time-budget system for 150 researchers, delivering 98% of owed GPU hours while maintaining 98% cluster occupancy.

When the Allen Institute for AI (Ai2) replaced its priority-based GPU scheduler with a system built on time budgets and hierarchical fair-share allocation, it reshaped how roughly 150 researchers compete for thousands of NVIDIA H100, B200, and B300 GPUs. The new design, rolled out cluster by cluster starting in late July, delivered 98% of owed GPU hours to research teams while holding cluster occupancy steady at 98%. The shift moved resource debates from case-by-case operational firefighting to a transparent budgeting process led by program managers.
The limits of priority queues
Ai2's infrastructure team manages clusters ranging from 88 to 1,024 GPUs, all built for large-scale distributed training across LLMs, VLMs, robotics reinforcement learning, and scientific agentic use cases. Demand runs 2 to 3 times higher than supply. Every available GPU hour has multiple competing workloads.
Under the old priority-based scheduler, researchers could opt out of preemptability and park non-preemptible workloads on their team's concurrent GPU limit indefinitely. The result was predictable. "Squatting" emerged as users left no-op jobs running so they could connect later without waiting in a queue. Priority inflation followed, with 100% of scheduled workloads eventually using HIGH priority and starving lower tiers. On-call engineers spent most of their ticket response time negotiating shutdowns of non-preemptible workloads on hosts with known maintenance problems.
"We had built a perfect laboratory for observing the 'tragedy of the commons,'" the Ai2 infrastructure team wrote. Individuals competing over a scarce shared resource achieved a non-optimal global result.
Budgets instead of monopolies
The team had already experimented with assigning GPU monopolies to important projects, but that left hardware idle when a team's research hit a seasonal lull. They needed the ownership incentive without sacrificing occupancy.
The answer was to allocate GPU time rather than physical GPUs. Managers distribute proportional time budgets across projects in a hierarchy that mirrors the research org chart. A leaf project, for example, holds a 35% claim on total cluster capacity regardless of activity elsewhere. Workloads that exceed their allocation can still run, but they are unprotected from preemption and are not charged to any budget.
This changed the incentive structure. "Nothing is free, so any trick to get GPU time draws from the benefiting user's allocation," the team explained. "A squatting workload is spending team budget on nothing. Our strategy is to make gaming the scheduler more expensive than honestly engaging in the debate for a larger budget."
The fair-share algorithm tracks occupancy over a sliding 7-day lookback window and sorts workloads from under-utilized allocations above those from over-utilized ones. It distinguishes allocated occupancy, which is charged to a budget and protected from preemption during a declared minimum runtime, from unallocated occupancy, which is free, unprotected, and used to soak up idle capacity.
A scheduling contract with time-slicing
Long-running training jobs - sometimes spanning days or weeks - made fair rebalancing nearly impossible under the old system. The new scheduling contract requires every workload to declare a minimum runtime. During that window the job is protected. Afterward, the scheduler can preempt and automatically re-queue resumable workloads to converge on fair-share targets.
This contract also automated host maintenance. Unhealthy nodes drain workloads as they reach their minimum runtime, eliminating the need for human negotiation. Repairs requiring a human in the loop fell by 74%.
Chris Clark, a researcher at Ai2, described the practical effect: "The new scheduler makes it feel like we have an extra 30% compute. In the old scheduler, if we had moments when we didn't need our full slot limit, that compute was basically lost. Now with the new scheduler, if that happens, we can later burst beyond our allocation limit and still see our jobs scheduled quickly and without preemption."
Simulation and rollout
Before deployment, the team built a simulation environment that tested scheduling decisions against both historical submission data and hand-crafted scenarios. One key hypothesis involved debug workloads - small jobs requiring 15 minutes or less of runtime. Simulations predicted p90 queue wait times for these jobs would drop from roughly 6 hours to 5 minutes.
Real-world results exceeded the forecast. Debug workload p90 wait time fell from 2 hours to 30 seconds. On the largest H100 cluster, median queue wait time dropped from 5 minutes to 24 seconds, and p90 wait time fell by roughly a third, from 2.8 hours to 1.8 hours.
Not every use case improved. Interactive sessions that researchers used for data analysis and code testing were previously held for up to a week. Under time-slicing, they faced an 8-hour protected runtime cap, after which they could be preempted, forcing researchers to rebuild volatile state by hand. Ai2 is responding with a CPU-only cluster for data prep sessions and plans to build restorable sessions to preserve the scheduling gains without the user friction.
Why this matters for IT and development professionals
The Ai2 case demonstrates that scheduler design is not just a queueing problem - it is an organizational design problem. Teams running shared GPU or CPU clusters can replace endless priority-tuning debates with a budgeting layer that forces explicit trade-offs. The combination of time budgets, fair-share algorithms, and a scheduling contract with minimum runtimes reduced on-call toil, eliminated squatting, and kept utilization high. For IT managers overseeing internal compute platforms, the lesson is that giving users ownership over time rather than hardware - and making preemption predictable rather than punitive - can align individual incentives with cluster-wide efficiency. Professionals looking to build these skills can explore AI Systems Admin Courses that cover resource management and scheduling for training infrastructure.