AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Ai2 describes replacing a priority-based GPU cluster scheduler with GPU time budgets, hierarchical fair-share allocation and a time-slicing contract. The organization says the approach shifts decisions about which research receives compute from day-to-day operational disputes to advance administrative budgeting; the report does not provide measured results for the new system.

Ai2 says it has replaced its priority-based GPU scheduler with a system built around GPU time budgets, hierarchical fair-share allocation and a time-slicing contract. The change is intended to direct scarce compute toward research leadership considers valuable while keeping clusters occupied, and to move arguments over access from live scheduling operations into an administrative budgeting process.

Ai2’s AI Infrastructure team says it manages thousands of NVIDIA H100, B200 and B300 GPUs across clusters ranging from 88 to 1,024 GPUs. The systems support about 150 internal researchers working on areas including large language and vision-language model training, robotics reinforcement-learning simulation and scientific agent development. According to the report, submitted workloads at any given time request two to three times the available GPU capacity.

Under the previous arrangement, workloads could be marked non-preemptible, with teams subject to limits on how many GPUs they could protect. Preemptible jobs could use otherwise idle capacity. Ai2 says the arrangement encouraged “GPU squatting”: users kept no-op workloads running so they could connect to a GPU quickly when needed. The team also says priority levels inflated until virtually all scheduled jobs were marked high priority, leaving lower-priority work starved of resources.

The new approach allocates portions of GPU time rather than ownership of particular GPUs. The report identifies three parts of the replacement: budgets set for GPU time, a hierarchical system for sharing resources and an agreement governing time-slicing. Ai2 says this lets leadership set relative research priorities before specific workloads arrive, while the scheduler allocates access as demand changes.

At a glance
reportWhen: Described in an Ai2 report published on…
The developmentAi2 has introduced a GPU scheduling system that allocates research projects time budgets and shared access rather than relying on workload priority alone.

Why Time Budgets Change Access

The change addresses a basic tension for research institutions: compute demand exceeds supply, but assigning fixed clusters to individual teams can leave hardware idle when those teams have no ready workloads. A time budget can preserve some ownership-like incentive—projects have a defined share to draw on—without tying capacity permanently to one group.

For researchers, the intended benefit is a clearer basis for access than competing to mark jobs high priority or negotiating directly with operators. For infrastructure staff, fewer protected workloads could also reduce the operational burden of asking users to stop jobs on machines that need maintenance. Those are goals described by Ai2, not independently verified outcomes: the report provides no before-and-after figures for occupancy, wait times, utilization, maintenance response or research output.

The design also changes where a difficult judgment is made. Rather than asking a scheduler to infer the scientific value of each job in real time, organizational leaders set budgets in advance. That can make trade-offs more explicit, but it does not remove disagreement over which projects deserve resources or guarantee that budgets reflect changing research needs.

Amazon

NVIDIA H100 GPU server

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Priority Queues to Budgets

Ai2 frames its infrastructure goals as a four-part hierarchy: hardware availability, workload occupancy, the impact of choosing which work runs, and GPU utilization within each workload. Its report focuses on the third measure—the impact of scheduling choices—rather than claiming that the new system has already improved every part of the hierarchy.

The earlier priority system produced behavior the team says was difficult to manage. Because users could opt out of preemption, engineers spent substantial ticket-response time negotiating shutdowns of protected jobs on hosts with maintenance issues. Attempts to correct the problems by tightening priority controls or granting important projects dedicated GPU monopolies did not resolve the underlying trade-off.

Dedicated allocations could strand hardware when one team’s research was between experiments while another team was waiting. Ai2 describes that mismatch as trying to fit shifting research demand into a static schedule. Its alternative is to decide how much GPU time projects should receive and let shared scheduling handle the day-to-day variation.

““Based on submitted workloads, at any moment in time we have outstanding requests for 2-3x more GPUs than are available.””

— Ai2 AI Infrastructure team, in the report

Amazon

GPU cluster management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Results and Allocation Rules Unreported

The supplied report excerpt explains the system’s design and the problems it is meant to address, but it does not provide quantitative results after deployment. It is unclear whether the new scheduler has reduced idle time, shortened wait times, improved GPU utilization or lowered the number of maintenance-related interventions.

The excerpt also does not spell out how budgets are calculated, how often they are reviewed, what happens when a project exhausts its allocation, or how the time-slicing contract handles urgent work and changing priorities. It is not clear how much flexibility researchers retain to use unclaimed capacity or how the system resolves disputes. Those details matter to judging whether the approach balances fairness, responsiveness and impact in practice.

Amazon

GPU time scheduling tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Track Deployment and Measured Outcomes

The next evidence to watch for is Ai2’s reporting on how the scheduler performs under real workloads. Useful measures would include GPU occupancy and utilization, job waiting times, preemption frequency and the time infrastructure staff spend resolving access and maintenance conflicts. The supplied material does not identify a future release date or a formal evaluation plan.

Further detail on budget-setting and time-slicing would also show how the system responds when demand shifts after allocations are set. Until those operating rules and outcome measures are published, the confirmed development is a change in allocation design—not proof that the new arrangement has increased research impact or eliminated the problems of the old one.

Amazon

high performance computing GPU

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What changed in Ai2’s GPU scheduler?

Ai2 says it moved from a priority-based system to GPU time budgets, hierarchical fair-share allocation and time-slicing, allocating time rather than giving teams fixed ownership of GPUs.

Why did Ai2 change its previous approach?

The team says the previous system encouraged GPU squatting and priority inflation, while non-preemptible jobs could complicate maintenance. Dedicated GPU allocations also risked leaving hardware idle when a team had no ready workload.

How much does demand exceed available GPU capacity?

Ai2 reports that submitted workloads request two to three times the GPUs available at any given moment. The report gives this as a current demand-to-capacity description, not a change over time.

Has the new scheduler been shown to improve utilization?

The supplied report describes the system and its goals but gives no before-and-after measurements. Whether it improves occupancy, utilization, waiting times or research outcomes remains unclear.

Source: rss

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Apple Wants Blacklisted Chinese RAM — and That Tells You How Bad the Squeeze Got

Apple is lobbying the US government to buy memory chips from Chinese firm CXMT, raising questions about supply shortages and national security concerns.

Show HN: Infinite canvas notes in the non-Euclidean Poincaré disk

A new project introduces an infinite, non-Euclidean canvas for note-taking using the Poincaré disk model, enabling unique visual organization.

Podman 6: machine usability improvements (2025)

Podman 6 introduces significant improvements in machine management, including provider-agnostic commands and simplified machine creation, enhancing user experience.

DeepSeek makes the V4 Pro price discount permanent

DeepSeek has announced that the discounted price for its V4 Pro model will become permanent, significantly reducing costs for users starting April 26, 2026.