Skip to main content
Back to Blog
Cloud ComputingAI/MLSystem Administration
6 October 202614 min readUpdated 9 October 2026

Impactful Scheduling for GPU Clusters

Impactful Scheduling for GPU Clusters Building a cluster scheduler that prioritizes high impact research while maintaining full occupancy Published October 9, 2026 A pyramid of...

By Hardware Team

Impactful Scheduling for GPU Clusters

Building a cluster scheduler that prioritizes high-impact research while maintaining full occupancy

Published October 9, 2026

A pyramid of GPU metrics

Ai2’s AI Infrastructure team provides GPU capacity for large, distributed training workloads. The team evaluates this work through four related metrics:

  1. Availability: How often hardware is healthy and ready for work.
  2. Occupancy: The share of available time assigned to a specific workload.
  3. Impact: How often the most valuable workloads are selected for resources.
  4. Utilization: The fraction of GPU capacity used during a workload’s lifetime.

This article focuses on improving the impact of scheduling decisions. Ai2 replaced a priority-based scheduler with a system built around GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. The change moved decisions about how much GPU time each research project should receive from case-by-case operational negotiations to a more transparent administrative budgeting process.

Overcommitting

Ai2 manages thousands of NVIDIA H100, B200, and B300 GPUs in clusters ranging from 88 to 1024 GPUs. These clusters support large-scale distributed AI model training for approximately 150 internal researchers working across areas such as LLM and VLM training, robotics reinforcement learning simulation, and post-training for scientific agentic applications.

Demand for GPU time substantially exceeds supply. At any given time, submitted workloads request two to three times more GPUs than are available. Each GPU hour therefore has two or three research workloads competing for it.

Historically, Ai2 used a priority-based scheduler and allowed workloads to opt out of preemption. Each team had a limit on the number of GPUs that could be used concurrently by workloads protected from preemption. Preemptible workloads could exceed that limit when GPUs were idle.

This approach produced several predictable problems:

  • GPU squatting: Users parked no-op workloads that they could connect to later because debugging workloads could not be launched with sufficiently low latency.
  • Priority inflation: Eventually, all scheduled workloads used HIGH priority, starving lower-priority work.
  • Operational overhead: On-call engineers spent much of their ticket-response time negotiating the shutdown of non-preemptible workloads on hosts requiring maintenance.

The tragedy of the commons

When these problems first appeared, the underlying causes were not immediately clear. Initial solutions focused on tighter control over priority settings and, eventually, on assigning GPU monopolies to important projects.

This created a version of the “tragedy of the commons.” Individuals competed for a scarce shared resource and, by maximizing their own outcomes, produced a less effective global result while abusing the resource itself.

Resource allocation combines algorithm design, economics, and system management. Users often understand the value of their own workloads better than an organization does, but they may also have incentives to conceal that value or retain resources even when doing so harms overall performance.

In their 2011 paper introducing Dominant Resource Fairness, Ghodsi et al. described a search company that provided dedicated machines only when users could guarantee high utilization. The company discovered that users added infinite loops to their code to artificially increase reported utilization. Although the hardware and workloads differ, the resource-allocation problems remain similar.

Budgets instead of fixed schedules

A conventional response to a tragedy of the commons is to privatize the shared resource, giving owners an incentive to maximize its value. Assigning teams exclusive GPU sets followed this principle, but it was too inflexible. Research demand is seasonal and varies across teams, so exclusive ownership could leave GPUs idle while another team waited for capacity.

The team was effectively solving a knapsack problem, trying to fit changing research needs into a static schedule. The goal was to preserve the incentive associated with ownership while keeping GPUs fully occupied.

Ai2 instead allocated portions of GPU time rather than assigning teams specific GPUs. Precise future demand cannot be forecast because it depends on the results of novel experiments. Strategic priorities, however, can be debated and established in advance. Managers can therefore decide how to fund each research effort with GPU time based on its expected impact, while the scheduler uses those decisions to prioritize arriving workloads.

The resulting hierarchical system allows managers to distribute GPU time proportionally among the projects and researchers they oversee. Project A1, for example, has a 35% claim on total capacity regardless of how many other projects are waiting elsewhere in the hierarchy.

Every GPU request must be funded by a budget or it remains unprotected from preemption. Under the former system, HIGH priority had no cost, and non-preemptibility allowed teams to occupy their concurrent GPU limits indefinitely. Under the new model, every protected request consumes the benefiting user’s allocation. A workload that occupies GPUs without doing useful work spends its team’s budget.

The goal is to make gaming the scheduler more expensive than honestly arguing for a larger allocation. Ai2 continues to refine its budget-review process, but it requires frequent opportunities for researchers to explain their needs and decisions by managers with sufficient context. Allocation decisions within a project are made by a lead researcher, within a program by a principal investigator, and across programs by a lead program manager or the CEO.

Hierarchical fair share

Ai2 paired GPU time budgets with a hierarchical fair-share scheduler that manages the actual occupancy of allocations across the program tree.

The algorithm is based on established approaches. Hierarchical fair share over a time window traces back to the Hadoop Fair Scheduler in 2009, and similar methods are used in SLURM’s Fair Tree and YARN’s Fair Scheduler. Ai2’s implementation uses a research-program hierarchy, with weights determined by manager-set budgets rather than static quotas.

The scheduler tracks occupancy over a sliding lookback window, set to seven days by default. It ranks workloads from underused allocations ahead of workloads from overused allocations. Over a week, each active group should receive its allocated GPU time if it continues submitting enough demand.

The scheduler distinguishes between two types of occupancy:

  • Allocated occupancy: Time charged to a budget. It affects fair-share calculations, and the workload is protected from preemption during its minimum runtime window.
  • Unallocated occupancy: Time that is not charged to a budget and is unprotected from the start. It can be preempted by any allocated request.

Unallocated occupancy allows the cluster to remain full when funded workloads are not ready to run. Teams do not need to decline otherwise-unused GPU cycles.

The scheduling contract

Distributed training makes fair allocation difficult because workloads can run for hours, days, or weeks. Once scheduled, a workload might retain its GPUs for a week or more, preventing other groups from receiving their budgeted time. Long-running workloads also made maintenance dependent on negotiations between on-call engineers and job owners.

Ai2 addressed this with a scheduling contract. When requesting cluster access, a workload must declare its minimum runtime, meaning the shortest period needed to make meaningful progress. The workload is protected from preemption during that period. Afterward, the scheduler can rebalance the cluster and automatically requeue workloads that can resume later.

A user can also set the minimum runtime to zero, making the GPU time unallocated. Such workloads are always preemptible, but they are not charged to a budget.

The workload lifecycle is:

  1. The workload is submitted with a minimum runtime and a declaration of whether it is resumable.
  2. The scheduler places it according to the fair-share algorithm, weighted by the ratio of actual occupancy to allocated time during the lookback window.
  3. The workload runs for its minimum runtime, which is charged to its allocation.
  4. It may continue running while its associated allocation continues to prioritize it. This additional time is also charged to the allocation.
  5. The workload may be preempted and requeued, returning to step two.
  6. The workload completes and releases its resources.

Together, these rules introduce time-slicing. Running workloads can be removed and requeued automatically, allowing fair share to converge and reducing the incentive to squat on GPUs. The rules also allow unhealthy hosts to drain workloads as they reach their minimum runtimes, making repairs more automated.

Repairs requiring a human in the loop fell by 74%, reducing on-call work substantially.

Simulation

Scheduling changes can have unintended consequences because the problem is zero-sum: giving time to one researcher takes time away from another. Users who lose access may seek new workarounds.

Before deploying the budget-based system, Ai2 built a simulation environment to predict where wait times might increase and to test configuration settings such as the lookback-window length and the maximum minimum runtime. The selected maximum minimum runtime was eight hours.

The simulator accepts workloads and their submission schedules, then models preemption and GPU assignment decisions. Given each workload’s requested GPU count and total runtime, it can jump between schedulable moments and analyze queue wait times, preemption events, and GPU-time distribution across projects over many simulated days within seconds. Ai2 tested both historical submission data and constructed scenarios.

One focus was the behavior of “debug workloads.” These jobs require a small number of GPUs and a minimum runtime of 15 minutes or less, enough to determine whether a job launches successfully or fails early because of a bug or configuration problem.

The team tested whether these workloads would wait less than large training jobs, which often require many GPUs and hours of runtime. Small jobs can fit into more scheduling opportunities, but the exact latency mattered. A wait of one or two minutes could enable interactive development, while a ten-minute wait could make the workflow impractical.

Because historical data did not contain enough debug-like workloads, the simulations used hand-built test cases. The results supported the hypothesis: p90 debug-workload wait times fell from approximately six hours to five minutes.

Results

Ai2 began a cluster-by-cluster rollout at the end of July. The evaluation focused on whether funded workloads received their allocated time, whether the cluster remained fully occupied, and whether researchers could understand the scheduler well enough to make informed decisions.

During the 30-day test period, teams received 98% of the GPU hours owed to them. Thirteen of 15 team allocations received at least 95%, while the lowest result was 90%. Cluster occupancy remained at 98% before and after the change, even though demand exceeded capacity by two to three times in both periods. Unallocated time represented 18% of delivered GPU time, helping maintain occupancy when funded workloads were not ready.

Real results exceeded the simulator’s directional predictions. On the new scheduler, p90 debug-workload queue time fell from two hours to 30 seconds, compared with the simulated reduction from six hours to five minutes. The smaller number of debug workloads in the baseline produced greater variance in those measurements.

Time-slicing also improved general queue latency. On the largest H100 cluster, median queue wait fell from five minutes to 24 seconds. P90 wait time fell by approximately one-third, from 2.8 hours to 1.8 hours.

The results addressed the three original problems:

  1. Squatting: Short debug workloads start in under a minute, reducing the value of keeping idle workloads running. The cost of squatting is charged to the squatter’s budget, reducing the allocation available when it is genuinely needed.
  2. Priority inflation: Workloads can still declare a priority, but it affects ordering only within a team. Managers have an incentive to monitor priorities across their group and use their budgets effectively.
  3. On-call toil: Unhealthy hosts drain automatically as workloads reach their minimum runtimes. Repairs requiring human intervention fell by 74%.

Challenges

The learning curve was steeper than expected. Because the rollout was incremental, researchers saw different behavior depending on the cluster they used. Interfaces also retained older terminology, including workload priority, even though some terms had changed meaning.

Documentation alone did not resolve the confusion. Live explanatory sessions were more effective because researchers could ask questions and the engineering team could explain the scheduler’s decisions, including the reasons behind them, using real examples.

These sessions helped move the organization away from frustration and informal theories toward broader communication about the GPU needs of individual experiments. Researchers began participating in budget discussions with a clearer understanding of the tradeoffs involved in accommodating new requests.

Ai2 also introduced visualizations showing how assigned GPU time compared with expected allocations and exposing the metric used to sort the workload queue. When a workload was preempted, users could consult these views to understand why. Budget owners could also see how GPU time was distributed among the projects they managed.

Not every use case improved. Researchers also use interactive sessions for data analysis and for testing training code during development. Under the old system, a researcher could keep such a session for up to a week. With time-slicing, sessions became subject to the eight-hour protected-runtime cap and could become preemptible when they exceeded their allocation.

Ai2 found that researchers depended more heavily than expected on the volatile state of these sessions. Preemption required both waiting for a new session and manually rebuilding the working state.

In response, Ai2 created two roadmap projects. The first is a CPU-only cluster adjacent to on-premises storage for development sessions focused on data preparation, preserving training-cluster capacity for workloads that require GPUs. The second is restorable sessions for these CPU-only workloads. This would allow sessions to be preempted for maintenance or time-slicing and restored elsewhere without requiring researchers to rebuild them manually.

Ai2 is also investigating capacity fragmentation, which could increase wait times for the largest workloads. Minimum-runtime protection may now apply to jobs that previously used preemptible mechanisms to exceed their team’s concurrent GPU limit. Those jobs could previously be interrupted at any time, which could waste work but also made it easier to assemble the GPUs needed by a large job.

The new scheduler may provide fewer opportunities to interrupt multiple workloads at once. Ai2 is using its simulator to reproduce this behavior while measuring production results.

The future

The next goal is utilization, the capstone of the GPU metrics model. Ai2 needs to make bootstrapping, checkpointing, and the training applications themselves as efficient as possible so that each workload gains the greatest possible value from its scheduled time.