Skip to main navigation Skip to search Skip to main content

Policy Optimization for Continuous Reinforcement Learning

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

We study reinforcement learning (RL) in the setting of continuous time and space, for an infinite horizon with a discounted objective and the underlying dynamics driven by a stochastic differential equation. Built upon recent advances in the continuous approach to RL, we develop a notion of occupation time (specifically for a discounted objective), and show how it can be effectively used to derive performance-difference and local-approximation formulas. We further extend these results to illustrate their applications in the PG (policy gradient) and TRPO/PPO (trust region policy optimization/proximal policy optimization) methods, which have been familiar and powerful tools in the discrete RL setting but under-developed in continuous RL. Through numerical experiments, we demonstrate the effectiveness and advantages of our approach. © 2023 Neural information processing systems foundation. All rights reserved.
Original languageEnglish
Title of host publication37th Conference on Neural Information Processing Systems (NeurIPS 2023)
EditorsA. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, S. Levine
PublisherNeural Information Processing Systems (NeurIPS)
Pages13637-13663
ISBN (Electronic)9781713899921
ISBN (Print)9781713899112
Publication statusPublished - Dec 2023
Externally publishedYes
Event37th Conference on Neural Information Processing Systems (NeurIPS 2023) - New Orleans Ernest N. Morial Convention Center, New Orleans, United States
Duration: 10 Dec 202316 Dec 2023
https://papers.nips.cc/paper_files/paper/2023
https://nips.cc/Conferences/2023

Publication series

NameAdvances in Neural Information Processing Systems
Volume36
ISSN (Print)1049-5258

Conference

Conference37th Conference on Neural Information Processing Systems (NeurIPS 2023)
Abbreviated titleNIPS '23
PlaceUnited States
CityNew Orleans
Period10/12/2316/12/23
Internet address

Fingerprint

Dive into the research topics of 'Policy Optimization for Continuous Reinforcement Learning'. Together they form a unique fingerprint.

Cite this