Skip to main navigation Skip to search Skip to main content

Retain to Refine: Adaptive Online Question Answering via Query Routing and Long-Short Memory

  • Yuchen Li (Co-first Author)
  • , Jiamin Chen (Co-first Author)
  • , Xinran Chen*
  • , Zhiyu Li
  • , Haojie Zhang*
  • , Rui Kong
  • , Jiayi Li
  • , Xinyu Ma
  • , Hengyi Cai
  • , Lixin Su
  • , Shuaiqiang Wang
  • , Jiashu Zhao
  • , Yongqi Zhang*
  • , Haoyi Xiong*
  • , Linghe Kong
  • , Lei Chen
  • , Dawei Yin*
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

4 Downloads (CityUHK Scholars)

Abstract

Large Language Models (LLMs) have shown strong capabilities in open-domain question answering (QA), but deploying them in real-world online systems introduces critical challenges. These include: (1) handling both simple and complex queries with appropriate levels of reasoning, (2) minimizing latency without compromising answer quality, and (3) maintaining answer consistency under evolving and noisy retrieval contexts. To address these challenges, we propose Retain-to-Refine (R2R), an adaptive agent-based QA framework designed for practical deployment. R2R integrates a Query Critic Agent (QCA) to assess query difficulty and route it accordingly: simple queries are answered directly using fast, prompt-based LLM calls, while complex queries are handled by a Memory Augmented Agent (MAA). MAA performs iterative reasoning guided by a unique long-short memory mechanism. Long-term memory retains and consolidates stable, core facts to ground the reasoning process, while short-term memory identifies transient information gaps to formulate highly focused subsequent queries. To ensure evidence quality, a Supervised Retrospection module validates and filters retrieved documents at each step. This agent-based design enables R2R to dynamically allocate computation based on question complexity, reducing unnecessary overhead while preserving high-quality answers when multi-step reasoning or external knowledge is required. Extensive evaluations across various settings and datasets demonstrate that the efficiency of R2R across diverse question types. In online settings, R2R delivers substantial gains in both response quality and efficiency, making it well-suited for large-scale industrial deployment in real-time QA services. © 2026 Copyright held by the owner/author(s).
Original languageEnglish
Title of host publicationKDD '26
Subtitle of host publicationProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1
PublisherAssociation for Computing Machinery
Pages2312-2322
Number of pages11
ISBN (Print)9798400722585
Publication statusPublished - 2026
Event32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026) - International Convention Center Jeju (ICC Jeju), Jeju Island, Korea, Republic of
Duration: 9 Aug 202613 Aug 2026
https://kdd2026.kdd.org/

Publication series

NameProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
ISSN (Print)2154-817X

Conference

Conference32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026)
Abbreviated titleACM KDD 2026
PlaceKorea, Republic of
CityJeju Island
Period9/08/2613/08/26
Internet address

Bibliographical note

Full text of this publication does not contain sufficient affiliation information. With consent from the author(s) concerned, the Research Unit(s) information for this record is based on the existing academic department affiliation of the author(s).

Research Keywords

  • agents
  • large language models
  • long-short memory
  • multi-hop question answering
  • query routing

Publisher's Copyright Statement

  • This full text is made available under CC-BY 4.0. https://creativecommons.org/licenses/by/4.0/

Fingerprint

Dive into the research topics of 'Retain to Refine: Adaptive Online Question Answering via Query Routing and Long-Short Memory'. Together they form a unique fingerprint.

Cite this