What type of study is this?

This is a Quantitative Study study (also classified as: Experimental Study).

September 30, 2025Open Access

Mutual Information Tracks Policy Coherence in Reinforcement Learning

Key Points

Successful learning shows characteristic mutual information growth from 0.84 to 2.83 bits, indicating selective attention in agents.
Joint mutual information follows an inverted U-curve, highlighting the shift from exploration to exploitation during training.
Differential diagnosis reveals that sensor faults broadly collapse information channels, while actuator faults disrupt action predictability.
Information patterns serve as signatures for learning and diagnostics, paving the way for adaptive RL systems able to detect faults.

Abstract

Reinforcement Learning (RL) agents deployed in real-world environments face degradation from sensor faults, actuator wear, and environmental shifts, yet lack intrinsic mechanisms to detect and diagnose these failures. We present an information-theoretic framework that reveals both the fundamental dynamics of RL and provides practical methods for diagnosing deployment-time anomalies. Through analysis of state-action mutual information patterns in a robotic control task, we first demonstrate that successful learning exhibits characteristic information signatures: mutual information between states and actions steadily increases from 0.84 to 2.83 bits (238% growth) despite growing state entropy, indicating that agents develop increasingly selective attention to task-relevant patterns. Intriguingly, states, actions and next states joint mutual information, MI(S,A;S'), follows an inverted U-curve, peaking during early learning before declining as the agent specializes suggesting a transition from broad exploration to efficient exploitation. More immediately actionable, we show that information metrics can differentially diagnose system failures: observation-space, i.e., states noise (sensor faults) produces broad collapses across all information channels with pronounced drops in state-action coupling, while action-space noise (actuator faults) selectively disrupts action-outcome predictability while preserving state-action relationships. This differential diagnostic capability demonstrated through controlled perturbation experiments enables precise fault localization without architectural modifications or performance degradation. By establishing information patterns as both signatures of learning and diagnostic for system health, we provide the foundation for adaptive RL systems capable of autonomous fault detection and policy adjustment based on information-theoretic principles.

Read Full Paperexternally

KI fragen

Bookmark

View Full Paper

Cite This Study

Reid et al. (Fri,) studied this question.

synapsesocial.com/papers/68dc1e358a7d58c25ebb19dd https://doi.org/https://doi.org/10.48550/arxiv.2509.10423

KI fragen

Bookmark

View Full Paper