Skip to main navigation Skip to search Skip to main content

STACKELBERG COUPLING OF ONLINE REPRESENTATION LEARNING AND REINFORCEMENT LEARNING

  • Fernando Martinez
  • , Tao Li*
  • , Yingdong Lu
  • , Juntao Chen
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

2 Downloads (CityUHK Scholars)

Abstract

Deep Q-learning jointly learns representations and values within monolithic networks, promising beneficial co-adaptation between features and value estimates. Although this architecture has attained substantial success, the coupling between representation and value learning creates instability as representations must constantly adapt to non-stationary value targets, while value estimates depend on these shifting representations. This is compounded by high variance in bootstrapped targets, which causes bias in value estimation in off-policy methods. We introduce Stackelberg Coupled Representation and Reinforcement Learning (SCORER), a framework for value-based RL that views representation and Q-learning as two strategic agents in a hierarchical game. SCORER models the Q-function as the leader, which commits to its strategy by updating less frequently, while the perception network (encoder) acts as the follower, adapting more frequently to learn representations that minimize Bellman error variance given the leader's committed strategy. Through this division of labor, the Q-function minimizes MSBE while perception minimizes its variance, thereby reducing bias accordingly, with asymmetric updates allowing stable co-adaptation, unlike simultaneous parameter updates in monolithic solutions. Our proposed SCORER framework leads to a bi-level optimization problem whose solution is approximated by a two-timescale algorithm that creates an asymmetric learning dynamic between the two players. Extensive experiments on DQN and its variants demonstrate that gains stem from algorithmic insight rather than model complexity.
Original languageEnglish
Title of host publication14th International Conference on Learning Representations (ICLR 2026)
EditorsC. Vondrick, B. Hariharan, C. Raffel, L. Pinto, D. Yang, A. Faust
PublisherInternational Conference on Learning Representations, ICLR
Publication statusPublished - 23 Apr 2026
Event14th International Conference on Learning Representations (ICLR 2026) - Riocentro Convention and Event Center, Rio de Janeiro, Brazil
Duration: 23 Apr 202627 Apr 2026
https://iclr.cc/Conferences/2026

Conference

Conference14th International Conference on Learning Representations (ICLR 2026)
Abbreviated titleICLR 2026
PlaceBrazil
CityRio de Janeiro
Period23/04/2627/04/26
Internet address

Bibliographical note

Research Unit(s) information for this publication is provided by the author(s) concerned.

Fingerprint

Dive into the research topics of 'STACKELBERG COUPLING OF ONLINE REPRESENTATION LEARNING AND REINFORCEMENT LEARNING'. Together they form a unique fingerprint.

Cite this