Skip to main navigation Skip to search Skip to main content

BOUNDED-PARAMETER PARTIALLY OBSERVABLE MARKOV DECISION PROCESSES: FRAMEWORK AND ALGORITHM

  • Yaodong NI
  • , Zhi-Qiang LIU

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

Abstract

Partially observable Markov decision processes (POMDPs) are powerful for planning under uncertainty. However, it is usually impractical to employ a POMDP with exact parameters to model the real-life situation precisely, due to various reasons such as limited data for learning the model, inability of exact POMDPs to model dynamic situations, etc. In this paper, assuming that the parameters of POMDPs are imprecise but bounded, we formulate the framework of bounded-parameter partially observable Markov decision processes (BPOMDPs). A modified value iteration is proposed as a basic strategy for tackling parameter imprecision in BPOMDPs. In addition, we design the UL-based value iteration algorithm, in which each value backup is based on two sets of vectors called U-set and L-set. We propose four strategies for computing U-set and L-set. We analyze theoretically the computational complexity and the reward loss of the algorithm. The effectiveness and robustness of the algorithm are shown empirically.

Original languageEnglish
Pages (from-to)821-863
Number of pages43
JournalInternational Journal of Uncertainty, Fuzziness and Knowlege-Based Systems
Volume21
Issue number6
DOIs
Publication statusPublished - Dec 2013

Research Keywords

  • Decision making under uncertainty
  • planning under uncertainty
  • bounded-parameter
  • POMDP
  • modified value iteration
  • ULVI algorithm
  • TRANSITION-PROBABILITIES
  • VALUE-ITERATION
  • APPROXIMATIONS
  • HORIZON
  • POMDPS
  • STATE

Fingerprint

Dive into the research topics of 'BOUNDED-PARAMETER PARTIALLY OBSERVABLE MARKOV DECISION PROCESSES: FRAMEWORK AND ALGORITHM'. Together they form a unique fingerprint.

Cite this