M3TQA: Massively Multilingual Multitask Table Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shu, Daixin, Yang, Jian, Wu, Zhenhe, Wu, Xianjie, Cheng, Xianfu, Guan, Xiangyuan, Wang, Yanghai, Wu, Pengfei, Yang, Tingyang, Zhu, Hualei, Zhang, Wei, Zhang, Ge, Liu, Jiaheng, Li, Zhoujun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908499070418944
author Shu, Daixin
Yang, Jian
Wu, Zhenhe
Wu, Xianjie
Cheng, Xianfu
Guan, Xiangyuan
Wang, Yanghai
Wu, Pengfei
Yang, Tingyang
Zhu, Hualei
Zhang, Wei
Zhang, Ge
Liu, Jiaheng
Li, Zhoujun
author_facet Shu, Daixin
Yang, Jian
Wu, Zhenhe
Wu, Xianjie
Cheng, Xianfu
Guan, Xiangyuan
Wang, Yanghai
Wu, Pengfei
Yang, Tingyang
Zhu, Hualei
Zhang, Wei
Zhang, Ge
Liu, Jiaheng
Li, Zhoujun
contents Tabular data is a fundamental component of real-world information systems, yet most research in table understanding remains confined to English, leaving multilingual comprehension significantly underexplored. Existing multilingual table benchmarks suffer from geolinguistic imbalance - overrepresenting certain languages and lacking sufficient scale for rigorous cross-lingual analysis. To address these limitations, we introduce a comprehensive framework for massively multilingual multitask table question answering, featuring m3TQA-Instruct, a large-scale benchmark spanning 97 languages across diverse language families, including underrepresented and low-resource languages. We construct m3TQA by curating 50 real-world tables in Chinese and English, then applying a robust six-step LLM-based translation pipeline powered by DeepSeek and GPT-4o, achieving high translation fidelity with a median BLEU score of 60.19 as validated through back-translation. The benchmark includes 2,916 professionally annotated question-answering pairs across four tasks designed to evaluate nuanced table reasoning capabilities. Experiments on state-of-the-art LLMs reveal critical insights into cross-lingual generalization, demonstrating that synthetically generated, unannotated QA data can significantly boost performance, particularly for low-resource languages. M3T-Bench establishes a new standard for multilingual table understanding, providing both a challenging evaluation platform and a scalable methodology for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16265
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle M3TQA: Massively Multilingual Multitask Table Question Answering
Shu, Daixin
Yang, Jian
Wu, Zhenhe
Wu, Xianjie
Cheng, Xianfu
Guan, Xiangyuan
Wang, Yanghai
Wu, Pengfei
Yang, Tingyang
Zhu, Hualei
Zhang, Wei
Zhang, Ge
Liu, Jiaheng
Li, Zhoujun
Computation and Language
Tabular data is a fundamental component of real-world information systems, yet most research in table understanding remains confined to English, leaving multilingual comprehension significantly underexplored. Existing multilingual table benchmarks suffer from geolinguistic imbalance - overrepresenting certain languages and lacking sufficient scale for rigorous cross-lingual analysis. To address these limitations, we introduce a comprehensive framework for massively multilingual multitask table question answering, featuring m3TQA-Instruct, a large-scale benchmark spanning 97 languages across diverse language families, including underrepresented and low-resource languages. We construct m3TQA by curating 50 real-world tables in Chinese and English, then applying a robust six-step LLM-based translation pipeline powered by DeepSeek and GPT-4o, achieving high translation fidelity with a median BLEU score of 60.19 as validated through back-translation. The benchmark includes 2,916 professionally annotated question-answering pairs across four tasks designed to evaluate nuanced table reasoning capabilities. Experiments on state-of-the-art LLMs reveal critical insights into cross-lingual generalization, demonstrating that synthetically generated, unannotated QA data can significantly boost performance, particularly for low-resource languages. M3T-Bench establishes a new standard for multilingual table understanding, providing both a challenging evaluation platform and a scalable methodology for future research.
title M3TQA: Massively Multilingual Multitask Table Question Answering
topic Computation and Language
url https://arxiv.org/abs/2508.16265