PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 10, 20250 citationsOpen Access

Semantic-Aware Scheduling for GPU Clusters with Large Language Models

View Full Paper
ZWZerui WangQHQinghao HuAKAna Klimovic

Key Points

  • SchedMate reduces average job completion times by up to 1.91x in deep learning workloads, improving overall efficiency.
  • The integration of three LLM-based components enhances existing deep learning schedulers while leveraging overlooked data sources.
  • Evaluations conducted on a 128-GPU physical cluster demonstrate reduced profiling overhead and better failure handling.
  • This innovative approach highlights the critical importance of semantic awareness in optimizing resource allocation for GPU clusters.

Abstract

Deep learning (DL) schedulers are pivotal in optimizing resource allocation in GPU clusters, but operate with a critical limitation: they are largely blind to the semantic context of the jobs they manage. This forces them to rely on limited metadata, leading to high profiling overhead, unreliable duration estimation, inadequate failure handling, and poor observability. To this end, we propose SchedMate, a framework that bridges this semantic gap by systematically extracting deep insights from overlooked, unstructured data sources: source code, runtime logs, and historical jobs. SchedMate enhances existing schedulers non-intrusively through three LLM-based components. Our implementation integrates seamlessly with existing deep learning schedulers. Evaluations on a 128-GPU physical cluster and extensive simulations on production traces show SchedMate reduces average job completion times by up to 1.91x, substantially enhancing the scheduling performance, demonstrating the critical role of semantic-awareness in modern DL scheduling.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2025) studied this question.

synapsesocial.com/papers/68e861b07ef2f04ca37e4bcbhttps://doi.org/10.48550/arxiv.2510.03334
Ask AI
Helpful
Bookmark
Share
View Full Paper