Skip to main navigation Skip to search Skip to main content

Joint MCS Adaptation and Beamforming Design for Multi-User MISO Systems: A Constrained Hybrid Deep Reinforcement Learning Approach

  • Xiaowen Ye
  • , Yuyi Mao
  • , Xianghao Yu*
  • , Liqun Fu
  • *Corresponding author for this work

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

Abstract

This article investigates the joint modulation-coding scheme (MCS) adaptation and beamforming design for multiuser multi-input single-output (MISO) systems, where one base station (BS) serves multiple user equipments (UEs) under imperfect and outdated channel state information (CSI). The sum-rate of the system is maximized while satisfying all UEs’ data rate requirements and the maximum transmit power constraint at the BS. Most existing beamforming designs overlooked that only a finite number of MCSs can be supported in practical communication systems. Moreover, previous works rely on perfect and real-time CSI for decision-making, neglecting processing delays and channel estimation errors. To circumvent the above issues, this article puts forth an intelligent joint optimization scheme based on deep reinforcement learning (DRL) techniques. Specifically, a new DRL framework, termed constrained hybrid DRL (CHDRL), is first proposed, which incorporates Lagrangian primal-dual optimization theory and a constrained action selection policy into conventional DRL to tackle various constraints. By integrating deep Q-network (DQN) and deep deterministic policy gradient (DDPG) algorithms, CHDRL is capable of simultaneously optimizing MCS in the discrete action domain and beamforming in the continuous action domain. In addition, to handle the large discrete action space of DQN, we develop an action branch architecture for CHDRL to enable independent and concurrent MCS decisions at different UEs. Finally, a multiparameter experience replay mechanism is designed to synchronously train Lagrangian multipliers and neural network parameters. The simulation results demonstrate that under imperfect and outdated CSI, CHDRL outperforms other benchmark schemes by: 1) achieving a significantly higher sum-rate; 2) meeting more UEs’ data rate requirements; and 3) being more robust against different CSI delays and the numbers of UEs. © 2014 IEEE.
Original languageEnglish
Pages (from-to)50852-50867
Number of pages16
JournalIEEE Internet of Things Journal
Volume12
Issue number23
Online published16 Sept 2025
DOIs
Publication statusPublished - 1 Dec 2025

Funding

The work of Xiaowen Ye was supported in part by the National Natural Science Foundation of China under Grant No. 62501157. The work of Liqun Fu was supported in part by the National Natural Science Foundation of China under Grant U23A20281 and in part by the National Social Science Foundation of China under Grant 24&ZD189.

Research Keywords

  • deep reinforcement learning
  • imperfect and outdated channel state information (CSI)
  • Joint modulation-coding scheme (MCS) adaptation and beamforming
  • Lagrangian primal-dual optimization
  • Deep reinforcement learning (DRL)

Fingerprint

Dive into the research topics of 'Joint MCS Adaptation and Beamforming Design for Multi-User MISO Systems: A Constrained Hybrid Deep Reinforcement Learning Approach'. Together they form a unique fingerprint.

Cite this