PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 20182,055 citationsOpen Access

Maximum Entropy Inverse Reinforcement Learning

View Full Paper
BZBrian D. ZiebartAMAndrew L. MaasJBJ. Andrew Bagnell

Key Points

  • To develop a probabilistic inverse reinforcement learning framework using the principle of maximum entropy to recover utility functions from noisy human demonstrations.
  • Formulated imitation learning within Markov Decision Processes by establishing a globally normalized probability distribution over decision sequences.
  • Applied the framework to real-world vehicle navigation datasets characterized by noisy and imperfect trajectory demonstrations.
  • Maintained theoretical performance guarantees matching standard inverse reinforcement learning algorithms while resolving ambiguity over demonstrated paths.
  • Successfully modeled human route preferences and enabled probabilistic inference of intended destinations and complete routes from partial trajectories.

Abstract

Recent research has shown the benefit of framing problems of imitation learning as solutions to Markov Decision Problems. This approach reduces learning to the problem of re- covering a utility function that makes the behavior induced by a near-optimal policy closely mimic demonstrated behavior. In this work, we develop a probabilistic approach based on the principle of maximum entropy. Our approach provides a well-defined, globally normalized distribution over decision sequences, while providing the same performance guarantees as existing methods. We develop our technique in the context of modeling real world navigation and driving behaviors where collected data is inherently noisy and imperfect. Our probabilistic approach enables modeling of route preferences as well as a powerful new approach to inferring destinations and routes based on partial trajectories.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ziebart et al. (2018) studied this question.

synapsesocial.com/papers/6a0a51af8e4d6c8168573ee1https://doi.org/10.1184/r1/6555512
Ask AI
Helpful
Bookmark
Share
View Full Paper