VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dang, Vy Tuong, Vo, An, Villa-Cueva, Emilio, Tau, Quang, Dm, Duc, Solorio, Thamar, Kim, Daeyoung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912842296328192
author Dang, Vy Tuong
Vo, An
Villa-Cueva, Emilio
Tau, Quang
Dm, Duc
Solorio, Thamar
Kim, Daeyoung
author_facet Dang, Vy Tuong
Vo, An
Villa-Cueva, Emilio
Tau, Quang
Dm, Duc
Solorio, Thamar
Kim, Daeyoung
contents We introduce VMMU, a Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark designed to evaluate how vision-language models (VLMs) interpret and reason over visual and textual information beyond English. VMMU consists of 2.5k multimodal questions across 7 tasks, covering a diverse range of problem contexts, including STEM problem solving, data interpretation, rule-governed visual reasoning, and abstract visual reasoning. All questions require genuine multimodal integration, rather than reliance on text-only cues or OCR-based shortcuts. We evaluate a diverse set of state-of-the-art proprietary and open-source VLMs on VMMU. Despite strong Vietnamese OCR performance, proprietary models achieve only 66% mean accuracy. Further analysis shows that the primary source of failure is not OCR, but instead multimodal grounding and reasoning over text and visual evidence. Code and data are available at https://vmmu-bench.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2508_13680
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
Dang, Vy Tuong
Vo, An
Villa-Cueva, Emilio
Tau, Quang
Dm, Duc
Solorio, Thamar
Kim, Daeyoung
Computation and Language
Machine Learning
We introduce VMMU, a Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark designed to evaluate how vision-language models (VLMs) interpret and reason over visual and textual information beyond English. VMMU consists of 2.5k multimodal questions across 7 tasks, covering a diverse range of problem contexts, including STEM problem solving, data interpretation, rule-governed visual reasoning, and abstract visual reasoning. All questions require genuine multimodal integration, rather than reliance on text-only cues or OCR-based shortcuts. We evaluate a diverse set of state-of-the-art proprietary and open-source VLMs on VMMU. Despite strong Vietnamese OCR performance, proprietary models achieve only 66% mean accuracy. Further analysis shows that the primary source of failure is not OCR, but instead multimodal grounding and reasoning over text and visual evidence. Code and data are available at https://vmmu-bench.github.io/
title VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2508.13680