Saved in:
Bibliographic Details
Main Authors: Guo, Jindi, Huang, Chaozheng, Fang, Xi
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.21277
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917436293382144
author Guo, Jindi
Huang, Chaozheng
Fang, Xi
author_facet Guo, Jindi
Huang, Chaozheng
Fang, Xi
contents We introduce MMTR-Bench, a benchmark designed to evaluate the intrinsic ability of Multimodal Large Language Models (MLLMs) to reconstruct masked text directly from visual context. Unlike conventional question-answering tasks, MMTR-Bench eliminates explicit prompts, requiring models to recover masked text from single- or multi-page inputs across real-world domains such as documents and webpages. This design isolates the reconstruction task from instruction-following abilities, enabling a direct assessment of a model's layout understanding, visual grounding, and knowledge integration. MMTR-Bench comprises 2,771 test samples spanning multiple languages and varying target lengths. To account for this diversity, we propose a level-aware evaluation protocol. Experiments on representative MLLMs show that the benchmark poses a significant challenge, especially for sentence- and paragraph-level reconstruction. The homepage is available at https://mmtr-bench-dataset.github.io/MMTR-Bench/.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21277
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Can MLLMs "Read" What is Missing?
Guo, Jindi
Huang, Chaozheng
Fang, Xi
Artificial Intelligence
We introduce MMTR-Bench, a benchmark designed to evaluate the intrinsic ability of Multimodal Large Language Models (MLLMs) to reconstruct masked text directly from visual context. Unlike conventional question-answering tasks, MMTR-Bench eliminates explicit prompts, requiring models to recover masked text from single- or multi-page inputs across real-world domains such as documents and webpages. This design isolates the reconstruction task from instruction-following abilities, enabling a direct assessment of a model's layout understanding, visual grounding, and knowledge integration. MMTR-Bench comprises 2,771 test samples spanning multiple languages and varying target lengths. To account for this diversity, we propose a level-aware evaluation protocol. Experiments on representative MLLMs show that the benchmark poses a significant challenge, especially for sentence- and paragraph-level reconstruction. The homepage is available at https://mmtr-bench-dataset.github.io/MMTR-Bench/.
title Can MLLMs "Read" What is Missing?
topic Artificial Intelligence
url https://arxiv.org/abs/2604.21277