DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhao, Yilun, Long, Yitao, Liu, Hongjun, Kamoi, Ryo, Nan, Linyong, Chen, Lyuhao, Liu, Yixin, Tang, Xiangru, Zhang, Rui, Cohan, Arman
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910561261846528
author Zhao, Yilun
Long, Yitao
Liu, Hongjun
Kamoi, Ryo
Nan, Linyong
Chen, Lyuhao
Liu, Yixin
Tang, Xiangru
Zhang, Rui
Cohan, Arman
author_facet Zhao, Yilun
Long, Yitao
Liu, Hongjun
Kamoi, Ryo
Nan, Linyong
Chen, Lyuhao
Liu, Yixin
Tang, Xiangru
Zhang, Rui
Cohan, Arman
contents Recent LLMs have demonstrated remarkable performance in solving exam-like math word problems. However, the degree to which these numerical reasoning skills are effective in real-world scenarios, particularly in expert domains, is still largely unexplored. This paper introduces DocMath-Eval, a comprehensive benchmark specifically designed to evaluate the numerical reasoning capabilities of LLMs in the context of understanding and analyzing specialized documents containing both text and tables. We conduct an extensive evaluation of 48 LLMs with Chain-of-Thought and Program-of-Thought prompting methods, aiming to comprehensively assess the capabilities and limitations of existing LLMs in DocMath-Eval. We found that even the current best-performing system (i.e., GPT-4o) still significantly lags behind human experts in solving complex numerical reasoning problems grounded in long contexts. We believe that DocMath-Eval can serve as a valuable benchmark for evaluating LLMs' capabilities in solving challenging numerical reasoning problems within expert domains.
format Preprint
id arxiv_https___arxiv_org_abs_2311_09805
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
Zhao, Yilun
Long, Yitao
Liu, Hongjun
Kamoi, Ryo
Nan, Linyong
Chen, Lyuhao
Liu, Yixin
Tang, Xiangru
Zhang, Rui
Cohan, Arman
Computation and Language
Recent LLMs have demonstrated remarkable performance in solving exam-like math word problems. However, the degree to which these numerical reasoning skills are effective in real-world scenarios, particularly in expert domains, is still largely unexplored. This paper introduces DocMath-Eval, a comprehensive benchmark specifically designed to evaluate the numerical reasoning capabilities of LLMs in the context of understanding and analyzing specialized documents containing both text and tables. We conduct an extensive evaluation of 48 LLMs with Chain-of-Thought and Program-of-Thought prompting methods, aiming to comprehensively assess the capabilities and limitations of existing LLMs in DocMath-Eval. We found that even the current best-performing system (i.e., GPT-4o) still significantly lags behind human experts in solving complex numerical reasoning problems grounded in long contexts. We believe that DocMath-Eval can serve as a valuable benchmark for evaluating LLMs' capabilities in solving challenging numerical reasoning problems within expert domains.
title DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
topic Computation and Language
url https://arxiv.org/abs/2311.09805