Los puntos clave no están disponibles para este artículo en este momento.
We study how representation learning can accelerate reinforcement learning rich observations, such as images, without relying either on domain or pixel-reconstruction. Our goal is to learn representations that provide for effective downstream control and invariance to task-irrelevant. Bisimulation metrics quantify behavioral similarity between states in MDPs, which we propose using to learn robust latent representations encode only the task-relevant information from observations. Our method encoders such that distances in latent space equal bisimulation in state space. We demonstrate the effectiveness of our method at task-irrelevant information using modified visual MuJoCo tasks, the background is replaced with moving distractors and natural videos, achieving SOTA performance. We also test a first-person highway driving where our method learns invariance to clouds, weather, and time of day. , we provide generalization results drawn from properties of metrics, and links to causal inference.
Zhang et al. (Thu,) studied this question.