PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 19, 20250 citationsOpen Access

Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters

View Full Paper
FSFoteini StratiZZZhendong ZhangGMGeorge Manos

Key Points

  • Sailor enhances training throughput while managing the diverse capabilities of GPU resources across zones.
  • The system significantly reduces stragglers by leveraging a dynamic search space exploration algorithm for job configurations.
  • By simulating runtime and memory footprints, Sailor optimizes distributed training on heterogeneous and geo-distributed clusters.
  • Supporting various types of heterogeneity, Sailor aims to make ML training on diverse resource pools more efficient.

Abstract

The high GPU demand of ML training makes it hard to allocate large homogeneous clusters of high-end GPUs in a single availability zone. Leveraging heterogeneous GPUs available within and across zones can improve throughput at a reasonable cost. However, training ML models on heterogeneous resources introduces significant challenges, such as stragglers and a large search space of possible job configurations. Current systems lack support for efficiently training models on heterogeneous resources. We present Sailor, a system that automates distributed training over heterogeneous, geo-distributed, and dynamically available resources. Sailor combines an efficient search space exploration algorithm, accurate runtime and memory footprint simulation, and a distributed training framework that supports different types of heterogeneity to optimize training throughput and cost.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Strati et al. (2025) studied this question.

synapsesocial.com/papers/68f43f92854d1061a58aca9bhttps://doi.org/10.48550/arxiv.2504.17096
Ask AI
Helpful
Bookmark
Share
View Full Paper