Skip to main navigation Skip to search Skip to main content

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

Abstract

Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models’ (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round retrieval and reasoning to resolve queries. Although recent agentic RAG has improved via reinforcement learning, they often incur substantial token overhead from search and reasoning. This trade-off prioritizes accuracy over efficiency. To address this issue, this work proposes TeaRAG, a Token-efficient agentic RAG framework capable of compressing both retrieval content and reasoning steps. 1) First, the retrieved content is compressed by augmenting chunk-based semantic retrieval with a graph retrieval using concise triplets. A knowledge association graph is then built from semantic similarity and co-occurrence. Finally, Personalized PageRank is leveraged to highlight key knowledge within this graph, reducing the number of tokens per retrieval. 2) Besides, to reduce reasoning steps, Iterative Process-aware Direct Preference Optimization (IP-DPO) is proposed. Specifically, our reward function evaluates the knowledge sufficiency by a knowledge matching mechanism, while penalizing excessive reasoning steps. This design can produce high-quality preference-pair datasets, supporting iterative DPO to improve reasoning conciseness. Across six datasets, TeaRAG improves the average Exact Match by 4% and 2% while reducing output tokens by 61% and 59% on Llama3‑8B‑Instruct and Qwen2.5‑14B‑Instruct, respectively. Code is available at https://github.com/Applied-Machine-Learning-Lab/TeaRAG. © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM.
Original languageEnglish
JournalACM Transactions on Information Systems
Online published13 Jun 2026
DOIs
Publication statusOnline published - 13 Jun 2026

Bibliographical note

Research Unit(s) information for this publication is provided by the author(s) concerned.

Funding

This work was supported in part by the grants from National Science and Technology Major Project (No. 2023ZD0121104), and the Anhui Natural Science Foundation (No. 2508085ZD006). Besides, this research was partially supported by National Natural Science Foundation of China (No.62502404), Hong Kong Research Grants Council (Research Impact Fund No.R1015-23, Collaborative Research Fund No.C1043-24GF, General Research Fund No. 11218325), Institute of Digital Medicine of City University of Hong Kong (No.9229503), Huawei (Huawei Innovation Research Program), Tencent (Tencent Rhino-Bird Focused Research Program, Tencent University Cooperation Project), Kuaishou (CCF-Kuaishou Large Model Explorer Fund No. 2025008, Kuaishou University Cooperation Project), Didi (CCF-Didi Gaia Scholars Research Fund), and Bytedance.

Research Keywords

  • Agentic Retrieval-Augmented Generation
  • Token-Efficient
  • Knowledge Graph
  • Process Supervision

Fingerprint

Dive into the research topics of 'TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework'. Together they form a unique fingerprint.

Cite this