Skip to main navigation Skip to search Skip to main content

TRIGO: Benchmarking Formal Mathematical Proof Reduction for Generative Language Models

  • Jing Xiong
  • , Jianhao Shen
  • , Ye Yuan
  • , Haiming Wang
  • , Yichun Yin
  • , Zhengying Liu
  • , Lin Li
  • , Zhijiang Guo
  • , Qingxing Cao
  • , Yinya Huang
  • , Chuanyang Zheng
  • , Xiaodan Liang*
  • , Ming Zhang*
  • , Qun Liu
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

55 Downloads (CityUHK Scholars)

Abstract

Automated theorem proving (ATP) has become an appealing domain for exploring the reasoning ability of the recent successful generative language models. However, current ATP benchmarks mainly focus on symbolic inference, but rarely involve the understanding of complex number combination reasoning. In this work, we propose TRIGO, an ATP benchmark that not only requires a model to reduce a trigonometric expression with step-by-step proofs but also evaluates a generative LM's reasoning ability on formulas and its capability to manipulate, group, and factor number terms. We gather trigonometric expressions and their reduced forms from the web, annotate the simplification process manually, and translate it into the “Lean” formal language system. We then automatically generate additional examples from the annotated samples to expand the dataset. Furthermore, we develop an automatic generator based on Lean-Gym to create dataset splits of varying difficulties and distributions in order to thoroughly analyze the model's generalization ability. Our extensive experiments show our proposed TRIGO poses a new challenge for advanced generative LM's including GPT-4 which is pre-trained on a considerable amount of open-source formal theorem-proving language data, and provide a new tool to study the generative LM's ability on both formal and mathematical reasoning. © 2023 Association for Computational Linguistics.
Original languageEnglish
Title of host publicationEMNLP 2023 - The 2023 Conference on Empirical Methods in Natural Language Processing
Subtitle of host publicationProceedings of the Conference
PublisherAssociation for Computational Linguistics
Pages11594-11632
ISBN (Print)9798891760608
DOIs
Publication statusPublished - Dec 2023
Event2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023) - Resorts World Convention Centre (Hybrid), Singapore
Duration: 6 Dec 202310 Dec 2023
https://aclanthology.org/2023.emnlp-main
https://2023.emnlp.org/

Publication series

NameEMNLP - Conference on Empirical Methods in Natural Language Processing, Proceedings

Conference

Conference2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023)
Abbreviated titleEMNLP
PlaceSingapore
Period6/12/2310/12/23
Internet address

Publisher's Copyright Statement

  • This full text is made available under CC-BY 4.0. https://creativecommons.org/licenses/by/4.0/

Fingerprint

Dive into the research topics of 'TRIGO: Benchmarking Formal Mathematical Proof Reduction for Generative Language Models'. Together they form a unique fingerprint.

Cite this