Skip to main navigation Skip to search Skip to main content

Hierarchize Pareto Dominance in Multi-Objective Stochastic Linear Bandits

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Multi-objective Stochastic Linear bandit (MOSLB) plays a critical role in the sequential decision-making paradigm, however, most existing methods focus on the Pareto dominance among different objectives without considering any priority. In this paper, we study bandit algorithms under mixed Pareto-lexicographic orders, which can reflect decision makers' preferences. We adopt the Grossone approach to deal with these orders and develop the notion of Pareto-lexicographic optimality to evaluate the learners' performance. Our work represents a first attempt to address these important and realistic orders in bandit algorithms. To design algorithms under these orders, the upper confidence bound (UCB) policy and the prior free lexicographical filter are adapted to approximate the optimal arms at each round. Moreover, the framework of the algorithms involves two stages in pursuit of the balance between exploration and exploitation. Theoretical analysis as well as numerical experiments demonstrate the effectiveness of our algorithms. © 2024, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved.
Original languageEnglish
Title of host publicationProceedings of the 38th Annual AAAI Conference on Artificial Intelligence
PublisherAssociation for the Advancement of Artificial Intelligence
Pages11489-11497
ISBN (Electronic)978-1-57735-887-9
DOIs
Publication statusPublished - 2024
Event38th Annual AAAI Conference on Artificial Intelligence (AAAI-24) - Vancouver Convention Centre – West Building, Vancouver, Canada
Duration: 20 Feb 202427 Feb 2024
https://aaai.org/aaai-conference/

Publication series

NameProceedings of the AAAI Conference on Artificial Intelligence
Number2
Volume38
ISSN (Print)2159-5399
ISSN (Electronic)2374-3468

Conference

Conference38th Annual AAAI Conference on Artificial Intelligence (AAAI-24)
Abbreviated titleAAAI-24
PlaceCanada
CityVancouver
Period20/02/2427/02/24
Internet address

Funding

The work described in this paper was supported by the Research Grants Council of the Hong Kong Special Administrative Region, China [GRF Project No: CityU 11215622] and by Natural Science Foundation of China (Project No: 62276223).

Research Keywords

  • ML: Reinforcement Learning
  • ML: Online Learning & Bandits

RGC Funding Information

  • RGC-funded

Fingerprint

Dive into the research topics of 'Hierarchize Pareto Dominance in Multi-Objective Stochastic Linear Bandits'. Together they form a unique fingerprint.

Cite this