PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 3, 2026Entropy0 citationsOpen Access

Information-Theoretic Intrinsic Motivation for Reinforcement Learning in Combinatorial Routing

RXRuozhang XiYNYao NiWWWangyu Wu

Key Points

  • Improved exploration efficiency was achieved through an information-theoretic framework in reinforcement learning.
  • Experimentation on the Travelling Salesman Problem and Split Delivery Vehicle Routing shows significant enhancement in solution quality.
  • The method uses neural mutual information estimators for better scalability in high-dimensional spaces without density modeling.
  • Intrinsic rewards are defined through mutual information, which reflects novelty in state-action transitions, enhancing training stability.

Abstract

Intrinsic motivation provides a principled mechanism for driving exploration in reinforcement learning when external rewards are sparse or delayed. A central challenge, however, lies in defining meaningful novelty signals in high-dimensional and combinatorial state spaces, where observation-level density estimation and prediction-error heuristics often become unreliable. In this work, we propose an information-theoretic framework for intrinsically motivated reinforcement learning grounded in the Information Bottleneck principle. Our approach learns compact latent state representations by explicitly balancing the compression of observations and the preservation of predictive information about future state transitions. Within this bottlenecked latent space, intrinsic rewards are defined through information-theoretic quantities that characterize the novelty of state-action transitions in terms of mutual information, rather than raw observation dissimilarity. To enable scalable estimation in continuous and high-dimensional settings, we employ neural mutual information estimators that avoid explicit density modeling and contrastive objectives based on the construction of positive-negative pairs. We evaluate the proposed method on two representative combinatorial routing problems, the Travelling Salesman Problem and the Split Delivery Vehicle Routing Problem, formulated as Markov decision processes with sparse terminal rewards. These problems serve as controlled testbeds for studying exploration and representation learning under long-horizon decision making. Experimental results demonstrate that the proposed information bottleneck-driven intrinsic motivation improves exploration efficiency, training stability, and solution quality compared to standard reinforcement learning baselines.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xi et al. (2026) studied this question.

synapsesocial.com/papers/69a75ad8c6e9836116a21331https://doi.org/10.3390/e28020140
Ask AI
Helpful
Bookmark
Share
View Full Paper