Abstract
Intrinsic motivation has emerged as a powerful approach to enhance exploration for Reinforcement Learning (RL). Intrinsically motivated RL provides intrinsic rewards to incentivize the agent to explore novel states. Despite its potential, current intrinsic motivation techniques face several drawbacks, such as static skill discovery, limited state coverage, sample inefficiency, and biases in policy optimization. Additionally, most research on intrinsically motivated RL focuses on general RL tasks, like controlling robots and playing video games, and it remains underexplored whether intrinsically motivated RL can improve the performance of adversarial attacks against RL.This thesis contributes to the field of RL by designing efficient intrinsic motivation methods for two main branches of RL---unsupervised RL and sparse-reward RL---and investigating the potential of intrinsically motivated RL in generating strong adversarial attacks against robotic RL agents and Large Language Models (LLMs). Firstly, for unsupervised RL, we introduce CIM, an efficient competence-based intrinsic motivation method that maximizes a novel lower bound of conditional state entropy. CIM significantly outperforms 15 baseline methods regarding performance (e.g., achieving 10 times larger state coverage) and sample efficiency (e.g., requiring 20 times fewer training samples) in various unsupervised RL environments. Secondly, for sparse-reward RL, we propose TUE, a simple yet efficient task-aware intrinsic motivation method that maximizes the state entropy in the task-aware state space instead of the original or latent state space. TUE significantly surpasses previous task-agnostic intrinsic motivation methods across various state-based and pixel-based sparse-reward RL tasks. Thirdly, for adversarial attacks against robotic RL agents, we develop IMAP, an adversarial policy learning method that utilizes intrinsically motivated RL to learn strong adversarial policies to uncover the potential vulnerabilities of the target robotic RL agent. We design four types of adversarial intrinsic regularizers for IMAP and empirically show that IMAP outperforms baseline methods across various single-agent and multi-agent RL environments. Lastly, for adversarial attacks against LLMs, we create MARIO, an automated red-teaming method that fine-tunes a red-team LLM via intrinsically motivated RL to generate stealthy and diverse adversarial prompts against the target LLM. MARIO shows superior performance in terms of diversity, quality, and stealthiness in text continuation and instruction following tasks.
In sum, this thesis presents two efficient intrinsic motivation methods (CIM and TUE) and two efficient intrinsically motivated RL-based adversarial attacks (IMAP and MARIO), demonstrates their significant performance improvements via comprehensive experiments, and highlights the importance of efficient intrinsically motivated RL in general and adversarial environments.
| Date of Award | 22 Aug 2024 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Cong WANG (Supervisor) |
Keywords
- Reinforcement Learning
- Intrinsic Motivation
- Adveresarial Attacks
- Large Language Models
Cite this
- Standard