Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gautam, Somraj, Penamakuri, Abhirama Subramanyam, Bhandari, Abhishek, Harit, Gaurav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909753252249600
author Gautam, Somraj
Penamakuri, Abhirama Subramanyam
Bhandari, Abhishek
Harit, Gaurav
author_facet Gautam, Somraj
Penamakuri, Abhirama Subramanyam
Bhandari, Abhishek
Harit, Gaurav
contents We introduce MMCRICBENCH-3K, a benchmark for Visual Question Answering (VQA) on cricket scorecards, designed to evaluate large vision-language models (LVLMs) on complex numerical and cross-lingual reasoning over semi-structured tabular images. MMCRICBENCH-3K comprises 1,463 synthetically generated scorecard images from ODI, T20, and Test formats, accompanied by 1,500 English QA pairs. It includes two subsets: MMCRICBENCH-E-1.5K, featuring English scorecards, and MMCRICBENCH-H-1.5K, containing visually similar Hindi scorecards, with all questions and answers kept in English to enable controlled cross-script evaluation. The task demands reasoning over structured numerical data, multi-image context, and implicit domain knowledge. Empirical results show that even state-of-the-art LVLMs, such as GPT-4o and Qwen2.5VL, struggle on the English subset despite it being their primary training language and exhibit a further drop in performance on the Hindi subset. This reveals key limitations in structure-aware visual text understanding, numerical reasoning, and cross-lingual generalization. The dataset is publicly available via Hugging Face at https://huggingface.co/datasets/DIALab/MMCricBench, to promote LVLM research in this direction.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17334
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs
Gautam, Somraj
Penamakuri, Abhirama Subramanyam
Bhandari, Abhishek
Harit, Gaurav
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
We introduce MMCRICBENCH-3K, a benchmark for Visual Question Answering (VQA) on cricket scorecards, designed to evaluate large vision-language models (LVLMs) on complex numerical and cross-lingual reasoning over semi-structured tabular images. MMCRICBENCH-3K comprises 1,463 synthetically generated scorecard images from ODI, T20, and Test formats, accompanied by 1,500 English QA pairs. It includes two subsets: MMCRICBENCH-E-1.5K, featuring English scorecards, and MMCRICBENCH-H-1.5K, containing visually similar Hindi scorecards, with all questions and answers kept in English to enable controlled cross-script evaluation. The task demands reasoning over structured numerical data, multi-image context, and implicit domain knowledge. Empirical results show that even state-of-the-art LVLMs, such as GPT-4o and Qwen2.5VL, struggle on the English subset despite it being their primary training language and exhibit a further drop in performance on the Hindi subset. This reveals key limitations in structure-aware visual text understanding, numerical reasoning, and cross-lingual generalization. The dataset is publicly available via Hugging Face at https://huggingface.co/datasets/DIALab/MMCricBench, to promote LVLM research in this direction.
title Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2508.17334