Abstract
We study reinforcement learning (RL) in the setting of continuous time and space, for an infinite horizon with a discounted objective and the underlying dynamics driven by a stochastic differential equation. Built upon recent advances in the continuous approach to RL, we develop a notion of occupation time (specifically for a discounted objective), and show how it can be effectively used to derive performance-difference and local-approximation formulas. We further extend these results to illustrate their applications in the PG (policy gradient) and TRPO/PPO (trust region policy optimization/proximal policy optimization) methods, which have been familiar and powerful tools in the discrete RL setting but under-developed in continuous RL. Through numerical experiments, we demonstrate the effectiveness and advantages of our approach. © 2023 Neural information processing systems foundation. All rights reserved.
| Original language | English |
|---|---|
| Title of host publication | 37th Conference on Neural Information Processing Systems (NeurIPS 2023) |
| Editors | A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, S. Levine |
| Publisher | Neural Information Processing Systems (NeurIPS) |
| Pages | 13637-13663 |
| ISBN (Electronic) | 9781713899921 |
| ISBN (Print) | 9781713899112 |
| Publication status | Published - Dec 2023 |
| Externally published | Yes |
| Event | 37th Conference on Neural Information Processing Systems (NeurIPS 2023) - New Orleans Ernest N. Morial Convention Center, New Orleans, United States Duration: 10 Dec 2023 → 16 Dec 2023 https://papers.nips.cc/paper_files/paper/2023 https://nips.cc/Conferences/2023 |
Publication series
| Name | Advances in Neural Information Processing Systems |
|---|---|
| Volume | 36 |
| ISSN (Print) | 1049-5258 |
Conference
| Conference | 37th Conference on Neural Information Processing Systems (NeurIPS 2023) |
|---|---|
| Abbreviated title | NIPS '23 |
| Place | United States |
| City | New Orleans |
| Period | 10/12/23 → 16/12/23 |
| Internet address |
Fingerprint
Dive into the research topics of 'Policy Optimization for Continuous Reinforcement Learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver