Abstract
The classical linear quadratic regulation (LQR) problem of linear systems by state feedback has been widely addressed. However, the LQR problem by dynamic output feedback with optimal transient performance remains open. The main reason is that the observer error inevitably leads to suboptimal transient performance of the closed-loop system. In this article, we propose an optimal dynamic output feedback learning control approach to solve the LQR problem of linear continuous-time systems with unknown dynamics. In particular, we propose a novel internal dynamics called the internal model. Unlike the classical p-copy internal model, it is driven by the input and output of the system, and the role of the proposed internal model is to compensate for the transient error of the observer such that the output feedback LQR problem is solved with guaranteed optimality. A model-free learning algorithm is developed to estimate the optimal control gain of the dynamic output feedback controller. The algorithm does not require any prior knowledge of the system matrices or the system's initial state, thus leading to an optimal solution to the model-free LQR problem. The effectiveness of the proposed method is illustrated using an aircraft control system.
| Original language | English |
|---|---|
| Pages (from-to) | 4124-4131 |
| Number of pages | 8 |
| Journal | IEEE Transactions on Automatic Control |
| Volume | 70 |
| Issue number | 6 |
| Early online date | 20 Jan 2025 |
| DOIs | |
| Publication status | Published - Jun 2025 |
| Externally published | Yes |
Bibliographical note
Publisher Copyright:© 2025 IEEE.
Funding
This work was supported in part by the National Key R&D Program of China under Grant 2021ZD0112600, in part by the National Natural Science Foundation of China under Grant 62373058, in part by the Beijing Natural Science Foundation under Grant L233003, in part by the Key Program of the National Natural Science Foundation of China under Grant 61933002, in part by the National Science Fund for Distinguished Young Scholars of China under Grant 62025301, in part by the Postdoctoral Fellowship Program of China Postdoctoral Science Foundation under Grant GZC20233407, and in part by the Basic Science Center Programs of National Natural Science Foundation of China under Grant 62088101.
Keywords
- Linear quadratic regulation (LQR)
- observer error
- optimal output feedback control
- reinforcement learning (RL)
Fingerprint
Dive into the research topics of 'Optimal Output Feedback Learning Control for Continuous-Time Linear Quadratic Regulation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver