RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Pengzuo, Yang, Yuhang, Zhu, Guangcheng, Ye, Chao, Gu, Hong, Lu, Xu, Xiao, Ruixuan, Bao, Bowen, He, Yijing, Zha, Liangyu, Ye, Wentao, Zhao, Junbo, Wang, Haobo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917145921716224
author Wu, Pengzuo
Yang, Yuhang
Zhu, Guangcheng
Ye, Chao
Gu, Hong
Lu, Xu
Xiao, Ruixuan
Bao, Bowen
He, Yijing
Zha, Liangyu
Ye, Wentao
Zhao, Junbo
Wang, Haobo
author_facet Wu, Pengzuo
Yang, Yuhang
Zhu, Guangcheng
Ye, Chao
Gu, Hong
Lu, Xu
Xiao, Ruixuan
Bao, Bowen
He, Yijing
Zha, Liangyu
Ye, Wentao
Zhao, Junbo
Wang, Haobo
contents With the rapid advancement of Large Language Models (LLMs), there is an increasing need for challenging benchmarks to evaluate their capabilities in handling complex tabular data. However, existing benchmarks are either based on outdated data setups or focus solely on simple, flat table structures. In this paper, we introduce RealHiTBench, a comprehensive benchmark designed to evaluate the performance of both LLMs and Multimodal LLMs (MLLMs) across a variety of input formats for complex tabular data, including LaTeX, HTML, and PNG. RealHiTBench also includes a diverse collection of tables with intricate structures, spanning a wide range of task types. Our experimental results, using 25 state-of-the-art LLMs, demonstrate that RealHiTBench is indeed a challenging benchmark. Moreover, we also develop TreeThinker, a tree-based pipeline that organizes hierarchical headers into a tree structure for enhanced tabular reasoning, validating the importance of improving LLMs' perception of table hierarchies. We hope that our work will inspire further research on tabular data reasoning and the development of more robust models. The code and data are available at https://github.com/cspzyy/RealHiTBench.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13405
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis
Wu, Pengzuo
Yang, Yuhang
Zhu, Guangcheng
Ye, Chao
Gu, Hong
Lu, Xu
Xiao, Ruixuan
Bao, Bowen
He, Yijing
Zha, Liangyu
Ye, Wentao
Zhao, Junbo
Wang, Haobo
Computation and Language
With the rapid advancement of Large Language Models (LLMs), there is an increasing need for challenging benchmarks to evaluate their capabilities in handling complex tabular data. However, existing benchmarks are either based on outdated data setups or focus solely on simple, flat table structures. In this paper, we introduce RealHiTBench, a comprehensive benchmark designed to evaluate the performance of both LLMs and Multimodal LLMs (MLLMs) across a variety of input formats for complex tabular data, including LaTeX, HTML, and PNG. RealHiTBench also includes a diverse collection of tables with intricate structures, spanning a wide range of task types. Our experimental results, using 25 state-of-the-art LLMs, demonstrate that RealHiTBench is indeed a challenging benchmark. Moreover, we also develop TreeThinker, a tree-based pipeline that organizes hierarchical headers into a tree structure for enhanced tabular reasoning, validating the importance of improving LLMs' perception of table hierarchies. We hope that our work will inspire further research on tabular data reasoning and the development of more robust models. The code and data are available at https://github.com/cspzyy/RealHiTBench.
title RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis
topic Computation and Language
url https://arxiv.org/abs/2506.13405