PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 16, 2026Journal of Engineering and Applied Science0 citationsOpen Access

A novel GPU cluster scheduling algorithm for cloud computing: a power-grid-inspired predictive-optimization framework

HLHai LiuYXYanling XiaoQZQian Zhou

Key Points

  • This research aims to improve the efficiency of GPU cluster scheduling in cloud computing by leveraging a power-grid-inspired framework.
  • Developed a hierarchical GPU scheduling algorithm based on economic dispatch and automatic generation control.
  • Implemented a prediction layer for future workload arrivals and a decision layer to prioritize tasks using a primal-dual scheme.
  • Validated the framework on an 8-GPU NVIDIA Tesla V100 cluster simulator.
  • Achieved a 30% improvement in mean GPU utilization.
  • Increased task throughput by 25%.
  • Reduced average job completion time by 27.5% across various workloads.

Abstract

Abstract Efficient management of GPU resources in cloud computing is critical for maximizing cluster throughput and meeting latency-sensitive service-level objectives (SLOs). GPU workloads exhibit bursty arrivals, heterogeneous resource profiles, and performance interference under co-location—properties analogous to stochastic generation and volatile loads in power grids. Static or purely reactive schedulers suffer from resource fragmentation and unstable utilization, similar to grid instability caused by uncontrolled intermittency. This paper introduces a hierarchical GPU cluster scheduling algorithm inspired by power grid economic dispatch and automatic generation control (AGC). At each epoch, the predictor layer estimates future arrivals and runtimes via state-space filtering, paralleling grid load forecasting. The decision layer computes multi-objective task priorities and solves constrained resource allocation through a primal–dual scheme, akin to optimal power flow with stability margins. The feedback layer adapts priority weights via online learning from observed delays and SLA violations, inspired by AGC’s continuous frequency regulation. Implemented on an 8-GPU NVIDIA Tesla V100 cluster simulator, the proposed method improves mean GPU utilization by up to 30%, task throughput by 25%, and reduces average job completion time by 27.5% across diverse workload regimes. Ablations confirm that prediction, optimization, and feedback are all essential for stable, high-efficiency GPU scheduling—demonstrating effective cross-domain transfer from power system operation to data center resource management.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2026) studied this question.

synapsesocial.com/papers/6a0809d7a487c87a6a40b9e3https://doi.org/10.1186/s44147-026-01033-3
Ask AI
Helpful
Bookmark
Share
View Full Paper