PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 18, 2026Machine Learning and Knowledge Extraction1 citationsOpen Access

Scenario-Guided Temporal Prototypes in Reinforcement Learning

View Full Paper
BDBlaž DobravecElektro Ljubljana (Slovenia)JŽJure ŽabkarUniversity of Ljubljana

Key Points

  • To develop an interpretable framework that improves decision-making in reinforcement learning by using recurring behavior patterns.
  • Introduced a framework for case-based decision making in reinforcement learning.
  • Grouped decision-making trajectories into recurring behavior patterns as prototypes.
  • Developed a local policy that links short-term patterns to actions using a similarity score.
  • The method provides pre hoc explanations for actions taken while maintaining high performance.
  • Clear explanations were generated for actions based on pattern recognition in simulations of CarRacing and voltage control.

Abstract

Deep reinforcement learning policies are hard to deploy in safety-critical settings, because they fail to explain why a sequence of actions is taken. We introduce an intrinsically interpretable framework that learns compact summaries of recurring behavior and uses them for case-based decision making. Our method (i) discovers global regimes by grouping trajectories into a small set of recurrent patterns and (ii) learns a prototype-conditioned local policy that maps the current short-horizon pattern to an action (“this matches prototype X → take action Y”). Each action is accompanied by a similarity score to relevant prototypes, which provide the explanations. We evaluate our approach on two domains: (1) CarRacing (pixel-based continuous control) and (2) a real voltage-control problem in low-voltage distribution networks. Our results indicate that the method provides clear pre hoc explanations while keeping task performance close to the reference policy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Dobravec et al. (2026) studied this question.

synapsesocial.com/papers/696c785beb60fb80d1396805https://doi.org/10.3390/make8010021
Ask AI
Helpful
Bookmark
Share
View Full Paper