Projects per year
Abstract
We propose a reinforcement learning (RL) approach to compute the expression of quasi-stationary distribution. Based on the fixed-point formulation of quasi-stationary distribution, we minimize the KL-divergence of two Markovian path distributions induced by candidate distribution and true target distribution. To solve this challenging minimization problem by gradient descent, we apply a reinforcement learning technique by introducing the reward and value functions. We derive the corresponding policy gradient theorem and design an actor-critic algorithm to learn the optimal solution and the value function. The numerical examples of finite state Markov chain are tested to demonstrate the new method.
| Original language | English |
|---|---|
| Article number | 133 |
| Journal | Entropy |
| Volume | 24 |
| Issue number | 1 |
| Online published | 17 Jan 2022 |
| DOIs | |
| Publication status | Published - Jan 2022 |
Funding
This research was funded by Government of Hong Kong, Grant Number 11305318; NSFC Grant Number 11871486.
Research Keywords
- Actor-critic algorithm
- KL-divergence
- Quasi-stationary distribution
- Reinforcement learning
Publisher's Copyright Statement
- This full text is made available under CC-BY 4.0. https://creativecommons.org/licenses/by/4.0/
RGC Funding Information
- RGC-funded
Fingerprint
Dive into the research topics of 'Learn Quasi-Stationary Distributions of Finite State Markov Chain'. Together they form a unique fingerprint.Projects
- 1 Finished
-
GRF: Explore Energy Landscape of Deep Learning to Understand Generalization Error
ZHOU, X. (Principal Investigator / Project Coordinator)
1/01/19 → 7/06/23
Project: Research
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver