Projects per year
Abstract
This paper studies the multiobjective bandit problem under lexicographic ordering, wherein the learner aims to simultaneously maximize m objectives hierarchically. The only existing algorithm for this problem considers the multi-armed bandit model, and its regret bound is O∼((KT)2/3) under a metric called priority-based regret. However, this bound is suboptimal, as the lower bound for single objective multi-armed bandits is Ω(K logT). Moreover, this bound becomes vacuous when the arm number K is infinite. To address these limitations, we investigate the multiobjective Lipschitz bandit model, which allows for an infinite arm set. Utilizing a newly designed multi-stage decision-making strategy, we develop an improved algorithm that achieves a general regret bound of O∼(T(d z i+1)/(d z i+2)) for the i-th objective, where dzi is the zooming dimension for the i-th objective, with i ∈ {1,2,...,m}. This bound matches the lower bound of the single objective Lipschitz bandit problem in terms of T, indicating that our algorithm is almost optimal. Numerical experiments confirm the effectiveness of our algorithm. © 2024, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved.
| Original language | English |
|---|---|
| Title of host publication | Proceedings of the 38th AAAI Conference on Artificial Intelligence |
| Editors | Jennifer Dy, Sriraam Natarajan, Michael Wooldridge |
| Place of Publication | Washington, DC |
| Publisher | AAAI Press |
| Pages | 16238-16246 |
| ISBN (Print) | 978-1-57735-887-9, 1-57735-887-2 |
| DOIs | |
| Publication status | Published - 2024 |
| Event | 38th Association for the Advancement of Artificial Intelligence Conference on Artificial Intelligence (AAAI-24) - Vancouver Convention Center, Vancouver, Canada Duration: 20 Feb 2024 → 27 Feb 2024 https://aaai.org/aaai-conference/ https://ojs.aaai.org/index.php/AAAI/issue/archive |
Publication series
| Name | Proceedings of the AAAI Conference on Artificial Intelligence |
|---|---|
| Number | 15 |
| Volume | 38 |
| ISSN (Print) | 2159-5399 |
| ISSN (Electronic) | 2374-3468 |
Conference
| Conference | 38th Association for the Advancement of Artificial Intelligence Conference on Artificial Intelligence (AAAI-24) |
|---|---|
| Place | Canada |
| City | Vancouver |
| Period | 20/02/24 → 27/02/24 |
| Internet address |
Funding
The work described in this paper was supported by the Research Grants Council of the Hong Kong Special Administrative Region, China [GRF Project No. CityU 11215622] and by Natural Science Foundation of China [Project No: 62276223].
RGC Funding Information
- RGC-funded
Fingerprint
Dive into the research topics of 'Multiobjective Lipschitz Bandits under Lexicographic Ordering'. Together they form a unique fingerprint.Projects
- 1 Active
-
GRF: Few for Many: A Non-Pareto Approach for Many Objective Optimization
ZHANG, Q. (Principal Investigator / Project Coordinator)
1/01/23 → …
Project: Research
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver