Skip to main navigation Skip to search Skip to main content

Sim2Rec: A Simulator-based Decision-Making Approach to Optimize Real-World Long-term User Engagement in Sequential Recommender Systems

  • Xiong-Hui Chen
  • , Bowei He
  • , Yang Yu*
  • , Qingyang Li
  • , Zhiwei Qing
  • , Wenjie Shang
  • , Jieping Ye
  • , Chen Ma
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Long-term user engagement (LTE) optimization in sequential recommender systems (SRS) is shown to be suited by reinforcement learning (RL) which finds a policy to maximize long-term rewards. Meanwhile, RL has its shortcomings, particularly requiring a large number of online samples for exploration, which is risky in real-world applications. One of the appealing ways to avoid the risk is to build a simulator and learn the optimal recommendation policy in the simulator. In LTE optimization, the simulator is to simulate multiple users’ daily feedback for given recommendations. However, building a user simulator with no reality-gap, i.e., can predict user’s feedback exactly, is unrealistic because the users’ reaction patterns are complex and historical logs for each user are limited, which might mislead the simulator-based recommendation policy. In this paper, we present a practical simulator-based recommender policy training approach, Simulation-to-Recommendation (Sim2Rec) to handle the reality-gap problem for LTE optimization. Specifically, Sim2Rec introduces a simulator set to generate various possibilities of user behavior patterns, then trains an environment-parameter extractor to recognize users’ behavior patterns in the simulators. Finally, a context-aware policy is trained to make the optimal decisions on all of the variants of the users based on the inferred environment-parameters. The policy is transferable to unseen environments (e.g., the real world) directly as it has learned to recognize all various user behavior patterns and to make the correct decisions based on the inferred environment-parameters. Experiments are conducted in synthetic environments and a real-world large-scale ride-hailing platform, DidiChuxing. The results show that Sim2Rec achieves significant performance improvement, and produces robust recommendations in unseen environments. © 2023 IEEE.
Original languageEnglish
Title of host publicationProceedings - 2023 IEEE 39th International Conference on Data Engineering ICDE 2023
PublisherIEEE
Pages3389-3402
ISBN (Electronic)979-8-3503-2227-9
ISBN (Print)979-8-3503-2228-6
DOIs
Publication statusPublished - 2023
Event39th IEEE International Conference on Data Engineering (ICDE 2023) - Marriott Anaheim, Anaheim, United States
Duration: 3 Apr 20237 Apr 2023
https://icde2023.ics.uci.edu/

Publication series

NameInternational Conference on Data Engineering
ISSN (Print)1063-6382
ISSN (Electronic)2375-026X

Conference

Conference39th IEEE International Conference on Data Engineering (ICDE 2023)
Abbreviated titleIEEE ICDE 2023
PlaceUnited States
CityAnaheim
Period3/04/237/04/23
Internet address

Bibliographical note

Research Unit(s) information for this publication is provided by the author(s) concerned.

Funding

This work is supported by the National Key Research and Development Program of China (2020AAA0107200), the National Science Foundation of China (61921006) and the Major Key Project of PCL (PCL2021A12).

Research Keywords

  • reinforcement learning
  • reality gaps
  • recommender systems

Fingerprint

Dive into the research topics of 'Sim2Rec: A Simulator-based Decision-Making Approach to Optimize Real-World Long-term User Engagement in Sequential Recommender Systems'. Together they form a unique fingerprint.

Cite this