PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 31, 20260 citationsOpen Access

Low-Rank Reinforcement Learning With Heterogeneous Human Feedback: From Recommendation to Large Language Models

View Full Paper
SLSeong Jin Lee

Key Points

  • This research aims to develop and analyze low-rank reinforcement learning methods to improve decision-making in high-dimensional environments with heterogeneous human feedback.
  • Investigated dynamic assortment problem in high-dimensional e-commerce settings with low-rank user–item interactions.
  • Proposed low-rank contextual reinforcement learning from human feedback framework for large language models.
  • Provided theoretical analyses and extensive numerical experiments to assess performance.
  • Achieved significant reduction in complexity for estimating personalized utilities, improving statistical efficiency.
  • Demonstrated provable regret bounds that show gains in efficiency over traditional methods.
  • The proposed framework offered robust performance and theoretical guarantees on sample efficiency under distribution shifts.

Abstract

Modern decision-making systems, from online marketplaces to large language models (LLMs), increasingly rely on high-dimensional environments with human feedback. However, the inherent heterogeneity of user preferences and the massive scale of feature spaces pose significant challenges for statistical efficiency and robust alignment. This dissertation develops and analyzes low-rank reinforcement learning (RL) methods designed to exploit latent structures to achieve scalability and theoretical rigor. In the first part, we investigate the dynamic assortment problem in high-dimensional e-commerce settings. By imposing a low-rank structure on user–item interactions, we significantly reduce the complexity of estimating personalized utilities. We demonstrate how this structure enables efficient exploration-exploitation strategies and provide provable regret bounds that characterize the gain in efficiency over traditional methods. We then assess the performance of our method in the Expedia Hotel recommendation dataset. The second part of this dissertation extends these principles to Reinforcement Learning from Human Feedback (RLHF) within large-scale contextual environments. We propose a low-rank contextual RLHF framework that simultaneously addresses diverse user preferences and the intricate latent spaces typical of modern LLMs. Our approach incorporates personalized reward modeling for alignment, offering theoretical guarantees on sample efficiency and robust performance under distribution shifts. Throughout this work, we provide rigorous theoretical analyses, algorithmic descriptions, and extensive numerical experiments. Together, these contributions illustrate how a low-rank perspective unifies efficiency and robustness in personalized decision-making systems, providing a scalable path for aligning complex models with heterogeneous human values.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Seong Jin Lee (2026) studied this question.

synapsesocial.com/papers/6a1bd1555783ba022b6fcdc3https://doi.org/10.17615/cnch-zm82
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation2024 · 2 citations
  2. 2Robust Reinforcement Learning from Human Feedback for Large Language Models Fine-Tuning2025
  3. 3Leveraging Domain Knowledge for Efficient Reward Modelling in RLHF: A Case-Study in E-Commerce Opinion Summarization2024 · 1 citations
  4. 4Dynamic Rank Reinforcement Learning for Adaptive Low-Rank Multi-Head Self-Attention in Large Language Models2026
  5. 5Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning2024 · 1 citations