PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 26, 2026IEEE Transactions on Cybernetics2 citations

Output-Feedback Control of Linear Continuous-Time Systems Using Discounted Inverse Reinforcement Learning

View Full Paper
HWHan WuQHQinglei HuJZJianying Zheng

Key Points

  • The goal is to develop a control algorithm for continuous-time systems using output data while addressing challenges of partial observability.
  • Developed a state reconstruction method based on expert control data.
  • Introduced a model-free output-feedback DIRL algorithm to solve for unknown value functions.
  • Analyzed the convergence and uniqueness of the proposed algorithm.
  • The algorithm successfully recovers the expert control policy.
  • Demonstrated superior computational efficiency compared to state-of-the-art methods.

Abstract

This article proposes a novel discounted inverse reinforcement learning (DIRL) algorithm for linear quadratic (LQ) control of unknown continuous-time (CT) systems with partially observable states and an unknown discounted value function. Existing DIRL methods predominantly rely on full-state feedback, limiting their applicability to practical scenarios where only input-output data are available. To this end, a state reconstruction method is designed for the system controlled by an expert using the measured desired output. Based on this, a model-free output-feedback (OPFB) DIRL algorithm is presented to iteratively solve the unknown value function and the corresponding optimal OPFB control policy equivalent to the expert control policy. The convergence of the proposed algorithm and the nonuniqueness of solutions are rigorously analyzed. Finally, comprehensive simulations reveal the effectiveness of the proposed algorithm in recovering the expert control policy and its superior computational efficiency compared to state-of-the-art (SOTA) methods.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wu et al. (2026) studied this question.

synapsesocial.com/papers/6977032e722626c4468e840dhttps://doi.org/10.1109/tcyb.2026.3651519
Ask AI
Helpful
Bookmark
Share
View Full Paper