PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

Sequential Policy Gradient for Adaptive Hyperparameter Optimization

View Full Paper
ZLZheng LiJCJerry ChengHGHuanying Gu

Key Points

  • Models retrained with the proposed SPG method showed improved performance on original datasets, enhancing effectiveness.
  • Experiments reveal that SPG outperformed standard transfer fine-tuning across five diverse datasets, achieving notable gains.
  • The Sequential Policy Gradient method generates state-action trajectories in a single forward pass for optimization efficiency.
  • SPG demonstrates consistent performance gains of +0.2 to 7% across widely adopted models, indicating strong industrial applicability.

Abstract

Reinforcement learning is essential for neural architecture search and hyperparameter optimization, but the conventional approaches impede widespread use due to prohibitive time and computational costs. Inspired by DeepSeek-V3 multi-token prediction architecture, we propose Sequential Policy Gradient modeling (SPG), a novel trajectory generation paradigm for lightweight online hyperparameter optimization. In contrast to conventional policy gradient methods, SPG extends the base model with temporary modules, enabling it to generate state-action (padded) trajectories in a single forward pass. Our experiments demonstrate that models gain performance when retrained with SPG on their original datasets and also outperform standard transfer fine-tuning. We evaluate on five datasets spanning computer vision (ImageNet, COCO), natural language processing (GLUE, SQuAD), and audio (SUPERB) to assess the industrial applicability of SPG. The proposed method demonstrates consistent improvements across widely adopted models, achieving performance gains of +0. 27\%, with significantly low computational costs. Fully reproducible code and pre-trained models: https: //huggingface. co/UniversalAlgorithmic/SPG.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2025) studied this question.

synapsesocial.com/papers/68f6379bb481a140a36cf653https://doi.org/10.48550/arxiv.2506.15051
Ask AI
Helpful
Bookmark
Share
View Full Paper