Skip to main navigation Skip to search Skip to main content

PROXYQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models

  • Haochen Tan
  • , Zhijiang Guo
  • , Zhan Shi
  • , Lu Xu
  • , Zhili Liu
  • , Yunlong Feng
  • , Xiaoguang Li
  • , Yasheng Wang
  • , Lifeng Shang
  • , Qun Liu
  • , Linqi Song*
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

84 Downloads (CityUHK Scholars)

Abstract

Large Language Models (LLMs) have succeeded remarkably in understanding long-form contents. However, exploring their capability for generating long-form contents, such as reports and articles, has been relatively unexplored and inadequately assessed by existing benchmarks. The prevalent evaluation methods, which predominantly rely on crowdsourcing, are recognized for their labor-intensive nature and lack of efficiency, whereas automated metrics, such as the ROUGE score, demonstrate discordance with human judgment criteria. In this paper, we propose PROXYQA, an innovative framework dedicated to assessing long-text generation. PROXYQA comprises in-depth human-curated meta-questions spanning various domains, each accompanied by specific proxy-questions with pre-annotated answers. LLMs are tasked to generate extensive content in response to these meta-questions, by engaging an evaluator and incorporating the generated texts as contextual background, PROXYQA assesses the generated content's quality through the evaluator's accuracy in addressing the proxy-questions. We examine multiple LLMs, emphasizing PROXYQA's demanding nature as a high-quality assessment tool. Human evaluation demonstrates that the proxy-question method is notably self-consistent and aligns closely with human evaluative standards. The dataset and leaderboard is available at https://proxy-qa.com. © 2024 Association for Computational Linguistics.
Original languageEnglish
Title of host publicationThe 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024) - Proceedings of the Conference
EditorsLun-Wei Ku, Andre Martins, Vivek Srikumar
Place of PublicationKerrville, TX
PublisherAssociation for Computational Linguistics
Pages6806-6827
Number of pages22
Volume1: Long Papers
ISBN (Print)9798891760943
DOIs
Publication statusPublished - Aug 2024
Event62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024) - Centara Grand and Bangkok Convention Centre, Bangkok, Thailand
Duration: 11 Aug 202416 Aug 2024
https://aclanthology.org/2024.acl-long
https://2024.aclweb.org/
https://aclanthology.org/
https://aclanthology.org/2024.acl-tutorials
https://aclanthology.org/2024.findings-acl

Publication series

NameProceedings of the Annual Meeting of the Association for Computational Linguistics
ISSN (Print)0736-587X

Conference

Conference62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024)
Abbreviated titleACL2024
PlaceThailand
CityBangkok
Period11/08/2416/08/24
Internet address

Bibliographical note

Research Unit(s) information for this publication is provided by the author(s) concerned.

Funding

We would like to express our profound gratitude to the anonymous reviewers for their invaluable and insightful feedback. This research has been partially funded by the Research Grants Council of the Hong Kong SAR under Grant GRF 11217823 and the Collaborative Research Fund C1042-23GF. Additional support was provided by the National Natural Science Foundation of China under Grant 62371411, the InnoHK initiative, the Government of the HKSAR, and the Laboratory for AI-Powered Financial Technologies.

Publisher's Copyright Statement

  • This full text is made available under CC-BY 4.0. https://creativecommons.org/licenses/by/4.0/

RGC Funding Information

  • RGC-funded

Fingerprint

Dive into the research topics of 'PROXYQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models'. Together they form a unique fingerprint.

Cite this