MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912422920454144 |
|---|---|
| author | Tang, Jingqun Liu, Qi Ye, Yongjie Lu, Jinghui Wei, Shu Lin, Chunhui Li, Wanqing Mahmood, Mohamad Fitri Faiz Bin Feng, Hao Zhao, Zhen He, Yangfan Lu, Kuan Wang, Yanjie Liu, Yuliang Liu, Hao Bai, Xiang Huang, Can |
| author_facet | Tang, Jingqun Liu, Qi Ye, Yongjie Lu, Jinghui Wei, Shu Lin, Chunhui Li, Wanqing Mahmood, Mohamad Fitri Faiz Bin Feng, Hao Zhao, Zhen He, Yangfan Lu, Kuan Wang, Yanjie Liu, Yuliang Liu, Hao Bai, Xiang Huang, Can |
| contents | Text-Centric Visual Question Answering (TEC-VQA) in its proper format not only facilitates human-machine interaction in text-centric visual environments but also serves as a de facto gold proxy to evaluate AI models in the domain of text-centric scene understanding. Nonetheless, most existing TEC-VQA benchmarks have focused on high-resource languages like English and Chinese. Despite pioneering works to expand multilingual QA pairs in non-text-centric VQA datasets through translation engines, the translation-based protocol encounters a substantial "visual-textual misalignment" problem when applied to TEC-VQA. Specifically, it prioritizes the text in question-answer pairs while disregarding the visual text present in images. Moreover, it fails to address complexities related to nuanced meaning, contextual distortion, language bias, and question-type diversity. In this work, we tackle multilingual TEC-VQA by introducing MTVQA, the first benchmark featuring high-quality human expert annotations across 9 diverse languages, consisting of 6,778 question-answer pairs across 2,116 images. Further, by comprehensively evaluating numerous state-of-the-art Multimodal Large Language Models~(MLLMs), including Qwen2-VL, GPT-4o, GPT-4V, Claude3, and Gemini, on the MTVQA benchmark, it is evident that there is still a large room for performance improvement (Qwen2-VL scoring 30.9 versus 79.7 for human performance), underscoring the value of MTVQA. Additionally, we supply multilingual training data within the MTVQA dataset, demonstrating that straightforward fine-tuning with this data can substantially enhance multilingual TEC-VQA performance. We aspire that MTVQA will offer the research community fresh insights and stimulate further exploration in multilingual visual text comprehension. The project homepage is available at https://bytedance.github.io/MTVQA/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_11985 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering Tang, Jingqun Liu, Qi Ye, Yongjie Lu, Jinghui Wei, Shu Lin, Chunhui Li, Wanqing Mahmood, Mohamad Fitri Faiz Bin Feng, Hao Zhao, Zhen He, Yangfan Lu, Kuan Wang, Yanjie Liu, Yuliang Liu, Hao Bai, Xiang Huang, Can Computer Vision and Pattern Recognition Text-Centric Visual Question Answering (TEC-VQA) in its proper format not only facilitates human-machine interaction in text-centric visual environments but also serves as a de facto gold proxy to evaluate AI models in the domain of text-centric scene understanding. Nonetheless, most existing TEC-VQA benchmarks have focused on high-resource languages like English and Chinese. Despite pioneering works to expand multilingual QA pairs in non-text-centric VQA datasets through translation engines, the translation-based protocol encounters a substantial "visual-textual misalignment" problem when applied to TEC-VQA. Specifically, it prioritizes the text in question-answer pairs while disregarding the visual text present in images. Moreover, it fails to address complexities related to nuanced meaning, contextual distortion, language bias, and question-type diversity. In this work, we tackle multilingual TEC-VQA by introducing MTVQA, the first benchmark featuring high-quality human expert annotations across 9 diverse languages, consisting of 6,778 question-answer pairs across 2,116 images. Further, by comprehensively evaluating numerous state-of-the-art Multimodal Large Language Models~(MLLMs), including Qwen2-VL, GPT-4o, GPT-4V, Claude3, and Gemini, on the MTVQA benchmark, it is evident that there is still a large room for performance improvement (Qwen2-VL scoring 30.9 versus 79.7 for human performance), underscoring the value of MTVQA. Additionally, we supply multilingual training data within the MTVQA dataset, demonstrating that straightforward fine-tuning with this data can substantially enhance multilingual TEC-VQA performance. We aspire that MTVQA will offer the research community fresh insights and stimulate further exploration in multilingual visual text comprehension. The project homepage is available at https://bytedance.github.io/MTVQA/. |
| title | MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2405.11985 |