PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

Zero-Shot Reinforcement Learning Under Partial Observability

View Full Paper
SJScott JeenUniversity of CambridgeTBTom BewleyUniversity of BristolJCJonathan M. CullenUniversity of Cambridge

Key Points

  • Performance of zero-shot reinforcement learning methods declines under partial observability, impacting generalization.
  • In experiments, memory-based architectures outperformed memory-free baselines, showing much improved effectiveness.
  • Evaluation was conducted in environments with partially observable states, rewards, and changing dynamics.
  • Findings suggest that incorporating memory can mitigate the challenges posed by partial observability in RL.

Abstract

Recent work has shown that, under certain assumptions, zero-shot reinforcement learning (RL) methods can generalise to any unseen task in an environment after reward-free pre-training. Access to Markov states is one such assumption, yet, in many real-world applications, the Markov state is only partially observable. Here, we explore how the performance of standard zero-shot RL methods degrades when subjected to partially observability, and show that, as in single-task RL, memory-based architectures are an effective remedy. We evaluate our memory-based zero-shot RL methods in domains where the states, rewards and a change in dynamics are partially observed, and show improved performance over memory-free baselines. Our code is open-sourced via: https://enjeeneer.io/projects/bfms-with-memory/.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jeen et al. (2025) studied this question.

synapsesocial.com/papers/68f6379bb481a140a36cf7a7https://doi.org/10.48550/arxiv.2506.15446
Ask AI
Helpful
Bookmark
Share
View Full Paper