PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 10, 20250 citationsOpen Access

Wavelet Predictive Representations for Non-Stationary Reinforcement Learning

View Full Paper
MWMin WangXLXin LiYHYe He

Key Points

  • WISDOM enhances adaptability by leveraging wavelet analysis for non-stationary tasks, improving agent performance.
  • Key experiments reveal that WISDOM significantly surpasses existing NSRL baselines in sample efficiency.
  • The wavelet temporal difference update operator was theoretically proven to aid in tracking MDP evolution.
  • Multiscale wavelet coefficients capture both trends and variations, showcasing robust improvements in policy adaptation.

Abstract

The real world is inherently non-stationary, with ever-changing factors, such as weather conditions and traffic flows, making it challenging for agents to adapt to varying environmental dynamics. Non-Stationary Reinforcement Learning (NSRL) addresses this challenge by training agents to adapt rapidly to sequences of distinct Markov Decision Processes (MDPs). However, existing NSRL approaches often focus on tasks with regularly evolving patterns, leading to limited adaptability in highly dynamic settings. Inspired by the success of Wavelet analysis in time series modeling, specifically its ability to capture signal trends at multiple scales, we propose WISDOM to leverage wavelet-domain predictive task representations to enhance NSRL. WISDOM captures these multi-scale features in evolving MDP sequences by transforming task representation sequences into the wavelet domain, where wavelet coefficients represent both global trends and fine-grained variations of non-stationary changes. In addition to the auto-regressive modeling commonly employed in time series forecasting, we devise a wavelet temporal difference (TD) update operator to enhance tracking and prediction of MDP evolution. We theoretically prove the convergence of this operator and demonstrate policy improvement with wavelet task representations. Experiments on diverse benchmarks show that WISDOM significantly outperforms existing baselines in both sample efficiency and asymptotic performance, demonstrating its remarkable adaptability in complex environments characterized by non-stationary and stochastically evolving tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2025) studied this question.

synapsesocial.com/papers/68e997abe14057276da7f1cahttps://doi.org/10.48550/arxiv.2510.04507
Ask AI
Helpful
Bookmark
Share
View Full Paper