TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Ce, Liu, Xiaofan, Song, Zhiyan, Chi, Ce, Zhao, Chen, Yang, Jingjing, Wang, Zhendong, Yang, Kexin, Shi, Boshen, Wang, Xing, Deng, Chao, Feng, Junlan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912480584794112
author Li, Ce
Liu, Xiaofan
Song, Zhiyan
Chi, Ce
Zhao, Chen
Yang, Jingjing
Wang, Zhendong
Yang, Kexin
Shi, Boshen
Wang, Xing
Deng, Chao
Feng, Junlan
author_facet Li, Ce
Liu, Xiaofan
Song, Zhiyan
Chi, Ce
Zhao, Chen
Yang, Jingjing
Wang, Zhendong
Yang, Kexin
Shi, Boshen
Wang, Xing
Deng, Chao
Feng, Junlan
contents The majority of data in businesses and industries is stored in tables, databases, and data warehouses. Reasoning with table-structured data poses significant challenges for large language models (LLMs) due to its hidden semantics, inherent complexity, and structured nature. One of these challenges is lacking an effective evaluation benchmark fairly reflecting the performances of LLMs on broad table reasoning abilities. In this paper, we fill in this gap, presenting a comprehensive table reasoning evolution benchmark, TReB, which measures both shallow table understanding abilities and deep table reasoning abilities, a total of 26 sub-tasks. We construct a high quality dataset through an iterative data processing procedure. We create an evaluation framework to robustly measure table reasoning capabilities with three distinct inference modes, TCoT, PoT and ICoT. Further, we benchmark over 20 state-of-the-art LLMs using this frame work and prove its effectiveness. Experimental results reveal that existing LLMs still have significant room for improvement in addressing the complex and real world Table related tasks. Both the dataset and evaluation framework are publicly available, with the dataset hosted on huggingface.co/datasets/JT-LM/JIUTIAN-TReB and the framework on github.com/JT-LM/jiutian-treb.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18421
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
Li, Ce
Liu, Xiaofan
Song, Zhiyan
Chi, Ce
Zhao, Chen
Yang, Jingjing
Wang, Zhendong
Yang, Kexin
Shi, Boshen
Wang, Xing
Deng, Chao
Feng, Junlan
Computation and Language
Artificial Intelligence
The majority of data in businesses and industries is stored in tables, databases, and data warehouses. Reasoning with table-structured data poses significant challenges for large language models (LLMs) due to its hidden semantics, inherent complexity, and structured nature. One of these challenges is lacking an effective evaluation benchmark fairly reflecting the performances of LLMs on broad table reasoning abilities. In this paper, we fill in this gap, presenting a comprehensive table reasoning evolution benchmark, TReB, which measures both shallow table understanding abilities and deep table reasoning abilities, a total of 26 sub-tasks. We construct a high quality dataset through an iterative data processing procedure. We create an evaluation framework to robustly measure table reasoning capabilities with three distinct inference modes, TCoT, PoT and ICoT. Further, we benchmark over 20 state-of-the-art LLMs using this frame work and prove its effectiveness. Experimental results reveal that existing LLMs still have significant room for improvement in addressing the complex and real world Table related tasks. Both the dataset and evaluation framework are publicly available, with the dataset hosted on huggingface.co/datasets/JT-LM/JIUTIAN-TReB and the framework on github.com/JT-LM/jiutian-treb.
title TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.18421