Skip to main navigation Skip to search Skip to main content

Reinforcement Learning Control for a 2-DOF Helicopter With State Constraints: Theory and Experiments

  • Zhijia Zhao
  • , Weitian He
  • , Chaoxu Mu
  • , Tao Zou*
  • , Keum-Shik Hong
  • , Han-Xiong Li
  • *Corresponding author for this work

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

Abstract

This study focuses on the novel reinforcement learning control strategy of a nonlinear two-degrees-of-freedom (2-DOF) helicopter system for tracking the desired trajectory while minimizing the tracking error. First, gradient descent algorithm is incorporated in the context of the reinforcement learning control scheme to obtain the adaptive laws. Subsequently, considering the uncertainties in the nonlinear system, radial basis function (RBF) neural networks (NNs) are exploited to approximate the unknown internal dynamics. In contrast to the previous studies, aiming at accelerating the convergence in reinforcement learning control, a barrier Lyapunov function is constructed to constrain the states to ensure that the tracking error rapidly converges to a neighborhood of zero. Under the proposed control strategy, the states of the closed-loop system are proven to be semi-globally uniformly ultimately bounded through rigorous Lyapunov analyses, and the state constraints are satisfied. Furthermore, the simulations and experiments conducted on a Quanser laboratory platform reveal that the proposed control functions are suitable and effective. Note to Practitioners—This paper is motivated by designing a reinforcement learning control strategy to enhance online learning capability and control performance of the controller for a nonlinear 2-DOF helicopter system. The control framework is divided into the design of the critic and actor NNs, responsible primarily for evaluating the control performance and approximating uncertainties in the system separately. Unlike the adaptive NN control, the actor NN weights are updated by combining information of states and inputs from the critic NN. In addition, aiming at accelerating the convergence, a barrier Lyapunov function is constructed to constrain the states to ensure that the tracking error rapidly converges to a neighborhood of zero. Finally, the proposed control strategy is validated in simulation and experiment on the Quanser laboratory platform. © 2022 IEEE.
Original languageEnglish
Pages (from-to)157-167
Number of pages11
JournalIEEE Transactions on Automation Science and Engineering
Volume21
Issue number1
Online published28 Oct 2022
DOIs
Publication statusPublished - Jan 2024

Funding

This work was supported in part by the National Natural Science Foundation of China under Grant 62273112, Grant 52171331, and Grant 62022061; in part by the Scientific Research Projects of Guangzhou Education Bureau under Grant 202032793; in part by the Science and Technology Planning Project of Guangzhou City under Grant 202102010398, Grant 202102010411, and Grant 202201010758; in part by the Guangzhou University-Hong Kong University of Science and Technology Joint Research Collaboration Fund under Grant YH202205; in part by the Open Research Fund from the Guangdong Laboratory of Artificial Intelligence and Digital Economy [Shenzhen (SZ)] under Grant GML-KF-22-27; in part by the Tianjin Natural Science Foundation under Grant 20JCYBJC00880; and in part by the Korea Institute of Energy Technology Evaluation and Planning through the Auspices of the Ministry of Trade, Industry and Energy, Republic of Korea, under Grant 20213030020160.

Research Keywords

  • 2-DOF helicopter
  • Artificial neural networks
  • barrier Lyapunov function
  • Convergence
  • DC motors
  • gradient descent
  • Helicopters
  • Lyapunov methods
  • Propellers
  • Quanser laboratory platform
  • RBF neural networks
  • Reinforcement learning
  • Uncertainty

Fingerprint

Dive into the research topics of 'Reinforcement Learning Control for a 2-DOF Helicopter With State Constraints: Theory and Experiments'. Together they form a unique fingerprint.

Cite this