MMTU: A Massive Multi-Task Table Understanding and Reasoning Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xing, Junjie, He, Yeye, Zhou, Mengyu, Dong, Haoyu, Han, Shi, Chen, Lingjiao, Zhang, Dongmei, Chaudhuri, Surajit, Jagadish, H. V.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912952822530048
author Xing, Junjie
He, Yeye
Zhou, Mengyu
Dong, Haoyu
Han, Shi
Chen, Lingjiao
Zhang, Dongmei
Chaudhuri, Surajit
Jagadish, H. V.
author_facet Xing, Junjie
He, Yeye
Zhou, Mengyu
Dong, Haoyu
Han, Shi
Chen, Lingjiao
Zhang, Dongmei
Chaudhuri, Surajit
Jagadish, H. V.
contents Tables and table-based use cases play a crucial role in many important real-world applications, such as spreadsheets, databases, and computational notebooks, which traditionally require expert-level users like data engineers, data analysts, and database administrators to operate. Although LLMs have shown remarkable progress in working with tables (e.g., in spreadsheet and database copilot scenarios), comprehensive benchmarking of such capabilities remains limited. In contrast to an extensive and growing list of NLP benchmarks, evaluations of table-related tasks are scarce, and narrowly focus on tasks like NL-to-SQL and Table-QA, overlooking the broader spectrum of real-world tasks that professional users face. This gap limits our understanding and model progress in this important area. In this work, we introduce MMTU, a large-scale benchmark with over 28K questions across 25 real-world table tasks, designed to comprehensively evaluate models ability to understand, reason, and manipulate real tables at the expert-level. These tasks are drawn from decades' worth of computer science research on tabular data, with a focus on complex table tasks faced by professional users. We show that MMTU require a combination of skills -- including table understanding, reasoning, and coding -- that remain challenging for today's frontier models, where even frontier reasoning models like OpenAI GPT-5 and DeepSeek R1 score only around 69\% and 57\% respectively, suggesting significant room for improvement. We highlight key findings in our evaluation using MMTU and hope that this benchmark drives further advances in understanding and developing foundation models for structured data processing and analysis. Our code and data are available at https://github.com/MMTU-Benchmark/MMTU and https://huggingface.co/datasets/MMTU-benchmark/MMTU.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05587
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMTU: A Massive Multi-Task Table Understanding and Reasoning Benchmark
Xing, Junjie
He, Yeye
Zhou, Mengyu
Dong, Haoyu
Han, Shi
Chen, Lingjiao
Zhang, Dongmei
Chaudhuri, Surajit
Jagadish, H. V.
Artificial Intelligence
Computation and Language
Databases
Machine Learning
Tables and table-based use cases play a crucial role in many important real-world applications, such as spreadsheets, databases, and computational notebooks, which traditionally require expert-level users like data engineers, data analysts, and database administrators to operate. Although LLMs have shown remarkable progress in working with tables (e.g., in spreadsheet and database copilot scenarios), comprehensive benchmarking of such capabilities remains limited. In contrast to an extensive and growing list of NLP benchmarks, evaluations of table-related tasks are scarce, and narrowly focus on tasks like NL-to-SQL and Table-QA, overlooking the broader spectrum of real-world tasks that professional users face. This gap limits our understanding and model progress in this important area. In this work, we introduce MMTU, a large-scale benchmark with over 28K questions across 25 real-world table tasks, designed to comprehensively evaluate models ability to understand, reason, and manipulate real tables at the expert-level. These tasks are drawn from decades' worth of computer science research on tabular data, with a focus on complex table tasks faced by professional users. We show that MMTU require a combination of skills -- including table understanding, reasoning, and coding -- that remain challenging for today's frontier models, where even frontier reasoning models like OpenAI GPT-5 and DeepSeek R1 score only around 69\% and 57\% respectively, suggesting significant room for improvement. We highlight key findings in our evaluation using MMTU and hope that this benchmark drives further advances in understanding and developing foundation models for structured data processing and analysis. Our code and data are available at https://github.com/MMTU-Benchmark/MMTU and https://huggingface.co/datasets/MMTU-benchmark/MMTU.
title MMTU: A Massive Multi-Task Table Understanding and Reasoning Benchmark
topic Artificial Intelligence
Computation and Language
Databases
Machine Learning
url https://arxiv.org/abs/2506.05587