Projects per year
Abstract
Lifelong deep reinforcement learning (DRL) approaches are commonly employed to adapt continuously to new tasks without forgetting previously acquired knowledge. While current lifelong DRL methods have shown promising advancements in retaining acquired knowledge, they suffer from significant adaptation efforts (i.e., longer training duration) and suboptimal policy when transferring to a new task that significantly deviates from previously learned tasks, a phenomenon known as the few-shot generalization challenge. In this work, we propose a generic approach that equips existing lifelong DRL methods with the capability of few-shot generalization. First, we employ selective experience reuse by leveraging the experience of encountered states, improving adaptation training for new tasks. Then, a relaxed softmax function is applied to the target Q values to improve the accuracy of evaluated Q values, leading to more optimal policies. Finally, we measure and reduce the discrepancy in data distribution between the policy and off-policy samples, resulting in improved adaptation efficiency. Extensive experiments have been conducted on three typical benchmarks to compare our approach with six representative lifelong DRL methods and two state-of-the-art (SOTA) few-shot DRL methods regarding their training speed, episode return, and average return of all episodes. Experimental results substantiate that our method improves the return of six lifelong DRL methods by at least 25%.
© 2024 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission.
© 2024 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission.
| Original language | English |
|---|---|
| Pages (from-to) | 6843-6857 |
| Number of pages | 15 |
| Journal | IEEE Transactions on Neural Networks and Learning Systems |
| Volume | 36 |
| Issue number | 4 |
| Online published | 30 Apr 2024 |
| DOIs | |
| Publication status | Published - Apr 2025 |
Bibliographical note
Research Unit(s) information for this publication is provided by the author(s) concerned.Funding
This work was supported in part by the Hong Kong Research Grant Council through General Research Fund (GRF) under Grant 11218621 and in part by the Research Impact Fund (RIF) Project under Grant R5060-19
Research Keywords
- Few-shot generalization
- lifelong deep reinforcement learning (DRL)
- selective experience reuse
RGC Funding Information
- RGC-funded
Fingerprint
Dive into the research topics of 'Policy Correction and State-Conditioned Action Evaluation for Few-Shot Lifelong Deep Reinforcement Learning'. Together they form a unique fingerprint.-
RIF-ExtU-Lead: Edge Learning: the Enabling Technology for Distributed Big Data Analytics in Cloud-Edge Environment
Guo, S. (Main Project Coordinator [External]) & WANG, J. (Principal Investigator / Project Coordinator)
1/05/20 → …
Project: Research
-
GRF: Age of Information Centric Task Scheduling in Autonomous Driving Systems
WANG, J. (Principal Investigator / Project Coordinator) & Qiao, C. (Co-Investigator)
1/01/22 → 12/12/25
Project: Research
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver