VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912842296328192 |
|---|---|
| author | Dang, Vy Tuong Vo, An Villa-Cueva, Emilio Tau, Quang Dm, Duc Solorio, Thamar Kim, Daeyoung |
| author_facet | Dang, Vy Tuong Vo, An Villa-Cueva, Emilio Tau, Quang Dm, Duc Solorio, Thamar Kim, Daeyoung |
| contents | We introduce VMMU, a Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark designed to evaluate how vision-language models (VLMs) interpret and reason over visual and textual information beyond English. VMMU consists of 2.5k multimodal questions across 7 tasks, covering a diverse range of problem contexts, including STEM problem solving, data interpretation, rule-governed visual reasoning, and abstract visual reasoning. All questions require genuine multimodal integration, rather than reliance on text-only cues or OCR-based shortcuts. We evaluate a diverse set of state-of-the-art proprietary and open-source VLMs on VMMU. Despite strong Vietnamese OCR performance, proprietary models achieve only 66% mean accuracy. Further analysis shows that the primary source of failure is not OCR, but instead multimodal grounding and reasoning over text and visual evidence. Code and data are available at https://vmmu-bench.github.io/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_13680 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark Dang, Vy Tuong Vo, An Villa-Cueva, Emilio Tau, Quang Dm, Duc Solorio, Thamar Kim, Daeyoung Computation and Language Machine Learning We introduce VMMU, a Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark designed to evaluate how vision-language models (VLMs) interpret and reason over visual and textual information beyond English. VMMU consists of 2.5k multimodal questions across 7 tasks, covering a diverse range of problem contexts, including STEM problem solving, data interpretation, rule-governed visual reasoning, and abstract visual reasoning. All questions require genuine multimodal integration, rather than reliance on text-only cues or OCR-based shortcuts. We evaluate a diverse set of state-of-the-art proprietary and open-source VLMs on VMMU. Despite strong Vietnamese OCR performance, proprietary models achieve only 66% mean accuracy. Further analysis shows that the primary source of failure is not OCR, but instead multimodal grounding and reasoning over text and visual evidence. Code and data are available at https://vmmu-bench.github.io/ |
| title | VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2508.13680 |