Abstract
Large Language Models (LLMs) have shown strong capabilities in open-domain question answering (QA), but deploying them in real-world online systems introduces critical challenges. These include: (1) handling both simple and complex queries with appropriate levels of reasoning, (2) minimizing latency without compromising answer quality, and (3) maintaining answer consistency under evolving and noisy retrieval contexts. To address these challenges, we propose Retain-to-Refine (R2R), an adaptive agent-based QA framework designed for practical deployment. R2R integrates a Query Critic Agent (QCA) to assess query difficulty and route it accordingly: simple queries are answered directly using fast, prompt-based LLM calls, while complex queries are handled by a Memory Augmented Agent (MAA). MAA performs iterative reasoning guided by a unique long-short memory mechanism. Long-term memory retains and consolidates stable, core facts to ground the reasoning process, while short-term memory identifies transient information gaps to formulate highly focused subsequent queries. To ensure evidence quality, a Supervised Retrospection module validates and filters retrieved documents at each step. This agent-based design enables R2R to dynamically allocate computation based on question complexity, reducing unnecessary overhead while preserving high-quality answers when multi-step reasoning or external knowledge is required. Extensive evaluations across various settings and datasets demonstrate that the efficiency of R2R across diverse question types. In online settings, R2R delivers substantial gains in both response quality and efficiency, making it well-suited for large-scale industrial deployment in real-time QA services. © 2026 Copyright held by the owner/author(s).
| Original language | English |
|---|---|
| Title of host publication | KDD '26 |
| Subtitle of host publication | Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1 |
| Publisher | Association for Computing Machinery |
| Pages | 2312-2322 |
| Number of pages | 11 |
| ISBN (Print) | 9798400722585 |
| Publication status | Published - 2026 |
| Event | 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026) - International Convention Center Jeju (ICC Jeju), Jeju Island, Korea, Republic of Duration: 9 Aug 2026 → 13 Aug 2026 https://kdd2026.kdd.org/ |
Publication series
| Name | Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining |
|---|---|
| ISSN (Print) | 2154-817X |
Conference
| Conference | 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026) |
|---|---|
| Abbreviated title | ACM KDD 2026 |
| Place | Korea, Republic of |
| City | Jeju Island |
| Period | 9/08/26 → 13/08/26 |
| Internet address |
Bibliographical note
Full text of this publication does not contain sufficient affiliation information. With consent from the author(s) concerned, the Research Unit(s) information for this record is based on the existing academic department affiliation of the author(s).Research Keywords
- agents
- large language models
- long-short memory
- multi-hop question answering
- query routing
Publisher's Copyright Statement
- This full text is made available under CC-BY 4.0. https://creativecommons.org/licenses/by/4.0/
Fingerprint
Dive into the research topics of 'Retain to Refine: Adaptive Online Question Answering via Query Routing and Long-Short Memory'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver