PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

Mapping representations in Reinforcement Learning via Semantic Alignment for Zero-Shot Stitching

View Full Paper
ARAntonio RicciardiVMValentino MaiorcaLMLuca Moschella

Key Points

  • Our framework supports high performance after mapping embeddings from different agents, enhancing reusability.
  • The approach employs semantic alignment to learn transformations that maintain accuracy across tasks and variations.
  • Empirical results show effective zero-shot stitching in the CarRacing environment with changing backgrounds.
  • By enabling modular re-assembly, this work emphasizes robust reinforcement learning in dynamic settings.

Abstract

Deep Reinforcement Learning (RL) models often fail to generalize when even small changes occur in the environment's observations or task requirements. Addressing these shifts typically requires costly retraining, limiting the reusability of learned policies. In this paper, we build on recent work in semantic alignment to propose a zero-shot method for mapping between latent spaces across different agents trained on different visual and task variations. Specifically, we learn a transformation that maps embeddings from one agent's encoder to another agent's encoder without further fine-tuning. Our approach relies on a small set of "anchor" observations that are semantically aligned, which we use to estimate an affine or orthogonal transform. Once the transformation is found, an existing controller trained for one domain can interpret embeddings from a different (existing) encoder in a zero-shot fashion, skipping additional trainings. We empirically demonstrate that our framework preserves high performance under visual and task domain shifts. We empirically demonstrate zero-shot stitching performance on the CarRacing environment with changing background and task. By allowing modular re-assembly of existing policies, it paves the way for more robust, compositional RL in dynamically changing environments.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ricciardi et al. (2025) studied this question.

synapsesocial.com/papers/68f5fcd68d54a28a75cf1ff8https://doi.org/10.48550/arxiv.2503.01881
Ask AI
Helpful
Bookmark
Share
View Full Paper