Key points are not available for this paper at this time.
Model-based reinforcement learning (RL) enjoys several benefits, such as-efficiency and planning, by learning a model of the environment's. However, learning a global model that can generalize across different is a challenging task. To tackle this problem, we decompose the task learning a global dynamics model into two stages: (a) learning a context vector that captures the local dynamics, then (b) predicting the next conditioned on it. In order to encode dynamics-specific information into context latent vector, we introduce a novel loss function that encourages context latent vector to be useful for predicting both forward and backward. The proposed method achieves superior generalization ability across simulated robotics and control tasks, compared to existing RL schemes.
Lee et al. (Thu,) studied this question.