PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 5, 20250 citationsOpen Access

Boosting Universal LLM Reward Design through Heuristic Reward Observation Space Evolution

View Full Paper
ZHZiwen HengZZZixu ZhaoTWTianhao Wu

Key Points

  • The proposed framework enhances reward design by evolving the observation space for large language models.
  • A state execution table tracks historical usage and success rates of environment states to improve exploration.
  • The framework reconciles user-provided task descriptions with expert-defined success criteria effectively.
  • Comprehensive evaluations on benchmark RL tasks validate the effectiveness and stability of the proposed approach.

Abstract

Large Language Models (LLMs) are emerging as promising tools for automated reinforcement learning (RL) reward design, owing to their robust capabilities in commonsense reasoning and code generation. By engaging in dialogues with RL agents, LLMs construct a Reward Observation Space (ROS) by selecting relevant environment states and defining their internal operations. However, existing frameworks have not effectively leveraged historical exploration data or manual task descriptions to iteratively evolve this space. In this paper, we propose a novel heuristic framework that enhances LLM-driven reward design by evolving the ROS through a table-based exploration caching mechanism and a text-code reconciliation strategy. Our framework introduces a state execution table, which tracks the historical usage and success rates of environment states, overcoming the Markovian constraint typically found in LLM dialogues and facilitating more effective exploration. Furthermore, we reconcile user-provided task descriptions with expert-defined success criteria using structured prompts, ensuring alignment in reward design objectives. Comprehensive evaluations on benchmark RL tasks demonstrate the effectiveness and stability of the proposed framework. Code and video demos are available at jingjjjjjie.github.io/LLM2Reward.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Heng et al. (2025) studied this question.

synapsesocial.com/papers/68e24e59d6d66a53c2472eaahttps://doi.org/10.48550/arxiv.2504.07596
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs2024
  2. 2Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning2025
  3. 3Generating and Evolving Reward Functions for Highway Driving with Large Language Models2024
  4. 4REvolve: Reward Evolution with Large Language Models using Human Feedback2024 · 2 citations
  5. 5Training Fast Robot Policies with Slow Foundation Models2024