PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 14, 2026The Computer Journal0 citations

MemOpt: a memory optimization method for deep learning model training based on dual intelligent reinforcement learning

View Full Paper
YZYan ZengCJChanghui JiangCHChengchuang Huang

Key Points

  • The aim is to enhance memory optimization during deep learning model training without sacrificing accuracy.
  • Proposes MemOpt, a novel memory optimization technique.
  • Employs dual-agent reinforcement learning for strategy selection.
  • Dynamically adjusts memory management based on model structure and device capabilities.
  • Increases maximum batch size for training by up to 8.7%.
  • Boosts training throughput by as much as 42.3% compared to existing methods.
  • Maintains model accuracy while reducing memory overhead.

Abstract

Abstract In recent years, with the rapid growth in the scale of datasets and neural network models, there has been a significant imbalance between the memory requirements during model training and the memory resources available on training devices. Existing memory optimization techniques like recomputation, memory swapping, and their adaptive combinations do not fully consider the structural information of the model and overlook the impact of application costs and timing on training efficiency. Addressing this issue, this paper proposes a memory optimization method called MemOpt, which uses dual-agent reinforcement learning to dynamically search for appropriate memory optimization strategies and execution timing based on model structure and device information. It can optimize memory without compromising accuracy while minimizing additional overhead. Experimental results show that the MemOpt method significantly increases the maximum batch size for model training by up to 8.7% and training throughput by up to 42.3% compared with baseline methods. In the future, this method may find better applications in large-scale neural networks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zeng et al. (2026) studied this question.

synapsesocial.com/papers/699011172ccff479cfe5777bhttps://doi.org/10.1093/comjnl/bxag014
Ask AI
Helpful
Bookmark
Share
View Full Paper