TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912480584794112 |
|---|---|
| author | Li, Ce Liu, Xiaofan Song, Zhiyan Chi, Ce Zhao, Chen Yang, Jingjing Wang, Zhendong Yang, Kexin Shi, Boshen Wang, Xing Deng, Chao Feng, Junlan |
| author_facet | Li, Ce Liu, Xiaofan Song, Zhiyan Chi, Ce Zhao, Chen Yang, Jingjing Wang, Zhendong Yang, Kexin Shi, Boshen Wang, Xing Deng, Chao Feng, Junlan |
| contents | The majority of data in businesses and industries is stored in tables, databases, and data warehouses. Reasoning with table-structured data poses significant challenges for large language models (LLMs) due to its hidden semantics, inherent complexity, and structured nature. One of these challenges is lacking an effective evaluation benchmark fairly reflecting the performances of LLMs on broad table reasoning abilities. In this paper, we fill in this gap, presenting a comprehensive table reasoning evolution benchmark, TReB, which measures both shallow table understanding abilities and deep table reasoning abilities, a total of 26 sub-tasks. We construct a high quality dataset through an iterative data processing procedure. We create an evaluation framework to robustly measure table reasoning capabilities with three distinct inference modes, TCoT, PoT and ICoT. Further, we benchmark over 20 state-of-the-art LLMs using this frame work and prove its effectiveness. Experimental results reveal that existing LLMs still have significant room for improvement in addressing the complex and real world Table related tasks. Both the dataset and evaluation framework are publicly available, with the dataset hosted on huggingface.co/datasets/JT-LM/JIUTIAN-TReB and the framework on github.com/JT-LM/jiutian-treb. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_18421 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Li, Ce Liu, Xiaofan Song, Zhiyan Chi, Ce Zhao, Chen Yang, Jingjing Wang, Zhendong Yang, Kexin Shi, Boshen Wang, Xing Deng, Chao Feng, Junlan Computation and Language Artificial Intelligence The majority of data in businesses and industries is stored in tables, databases, and data warehouses. Reasoning with table-structured data poses significant challenges for large language models (LLMs) due to its hidden semantics, inherent complexity, and structured nature. One of these challenges is lacking an effective evaluation benchmark fairly reflecting the performances of LLMs on broad table reasoning abilities. In this paper, we fill in this gap, presenting a comprehensive table reasoning evolution benchmark, TReB, which measures both shallow table understanding abilities and deep table reasoning abilities, a total of 26 sub-tasks. We construct a high quality dataset through an iterative data processing procedure. We create an evaluation framework to robustly measure table reasoning capabilities with three distinct inference modes, TCoT, PoT and ICoT. Further, we benchmark over 20 state-of-the-art LLMs using this frame work and prove its effectiveness. Experimental results reveal that existing LLMs still have significant room for improvement in addressing the complex and real world Table related tasks. Both the dataset and evaluation framework are publicly available, with the dataset hosted on huggingface.co/datasets/JT-LM/JIUTIAN-TReB and the framework on github.com/JT-LM/jiutian-treb. |
| title | TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2506.18421 |