PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 23, 202424 citationsOpen Access

MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

View Full Paper
ZJZiheng JiangHLHaibin LinYZYinmin Zhong

Key Points

  • MegaScale achieves a model utilization of 55.2% while training a 175 billion parameter language model, showcasing significant efficiency.
  • The system utilizes 12,288 GPUs, improving model FLOPs Utilization by 1.34 times compared to previous methods.
  • This system is designed to co-optimize algorithmic and system components, enhancing communication and stability during training processes efficiently and effectively with diagnosis tools in place to monitor tasks closely and address issues promptly, ensuring ongoing performance optimization throughout its operation over prolonged training periods and jobs across thousands of GPUs deployed in tandem for effective outcomes per training protocol and goals stated in the system's design architecture and objectives established accordingly, aiming to maintain functionality throughout the entire operation process sustainably and responsively during initialization and procedure transitions across data segments in task loads and handler deployments for large model implementations, according to descriptions provided in the detailed performance review framework that guides operational assessments and analytics for ongoing study efforts and application use cases moving forward.

Abstract

We present the design, implementation and engineering experience in building and deploying MegaScale, a production system for training large language models (LLMs) at the scale of more than 10,000 GPUs. Training LLMs at this scale brings unprecedented challenges to training efficiency and stability. We take a full-stack approach that co-designs the algorithmic and system components across model block and optimizer design, computation and communication overlapping, operator optimization, data pipeline, and network performance tuning. Maintaining high efficiency throughout the training process (i.e., stability) is an important consideration in production given the long extent of LLM training jobs. Many hard stability issues only emerge at large scale, and in-depth observability is the key to address them. We develop a set of diagnosis tools to monitor system components and events deep in the stack, identify root causes, and derive effective techniques to achieve fault tolerance and mitigate stragglers. MegaScale achieves 55.2% Model FLOPs Utilization (MFU) when training a 175B LLM model on 12,288 GPUs, improving the MFU by 1.34x compared to Megatron-LM. We share our operational experience in identifying and fixing failures and stragglers. We hope by articulating the problems and sharing our experience from a systems perspective, this work can inspire future LLM systems research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jiang et al. (2024) studied this question.

synapsesocial.com/papers/68e77c94b6db6435876f0f28https://doi.org/10.48550/arxiv.2402.15627
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1SimpleScale: Simplifying the Training of an LLM Model Using 1024 GPUs2025 · 1 citations
  2. 2Efficient large-scale language model training on GPU clusters using megatron-LM2021 · 38 citations
  3. 3Scaling Intelligence: Designing Data Centers for Next-Gen Language Models2025
  4. 4Optimizing Distributed Training on Frontier for Large Language Models2024 · 15 citations
  5. 5Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models2026