Skip to main navigation Skip to search Skip to main content

Learn Quasi-Stationary Distributions of Finite State Markov Chain

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

101 Downloads (CityUHK Scholars)

Abstract

We propose a reinforcement learning (RL) approach to compute the expression of quasi-stationary distribution. Based on the fixed-point formulation of quasi-stationary distribution, we minimize the KL-divergence of two Markovian path distributions induced by candidate distribution and true target distribution. To solve this challenging minimization problem by gradient descent, we apply a reinforcement learning technique by introducing the reward and value functions. We derive the corresponding policy gradient theorem and design an actor-critic algorithm to learn the optimal solution and the value function. The numerical examples of finite state Markov chain are tested to demonstrate the new method.
Original languageEnglish
Article number133
JournalEntropy
Volume24
Issue number1
Online published17 Jan 2022
DOIs
Publication statusPublished - Jan 2022

Funding

This research was funded by Government of Hong Kong, Grant Number 11305318; NSFC Grant Number 11871486.

Research Keywords

  • Actor-critic algorithm
  • KL-divergence
  • Quasi-stationary distribution
  • Reinforcement learning

Publisher's Copyright Statement

  • This full text is made available under CC-BY 4.0. https://creativecommons.org/licenses/by/4.0/

RGC Funding Information

  • RGC-funded

Fingerprint

Dive into the research topics of 'Learn Quasi-Stationary Distributions of Finite State Markov Chain'. Together they form a unique fingerprint.

Cite this