MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909833726263296 |
|---|---|
| author | Zheng, Weihua Liu, Zhengyuan Chakraborty, Tanmoy Xu, Weiwen Gao, Xiaoxue Tan, Bryan Chen Zhengyu Zou, Bowei Liu, Chang Hu, Yujia Xie, Xing Yi, Xiaoyuan Yao, Jing Wang, Chaojun Li, Long Liu, Rui Liu, Huiyao Inoue, Koji Sumida, Ryuichi Kawahara, Tatsuya Xu, Fan Ye, Lingyu Tian, Wei Kim, Dongjun Jung, Jimin Seo, Jaehyung Wangsajaya, Nadya Yuki Duc, Pham Minh Saxena, Ojasva Nandi, Palash Tao, Xiyan Karlina, Wiwik Luong, Tuan Vasan, Keertana Arun Lee, Roy Ka-Wei Chen, Nancy F. |
| author_facet | Zheng, Weihua Liu, Zhengyuan Chakraborty, Tanmoy Xu, Weiwen Gao, Xiaoxue Tan, Bryan Chen Zhengyu Zou, Bowei Liu, Chang Hu, Yujia Xie, Xing Yi, Xiaoyuan Yao, Jing Wang, Chaojun Li, Long Liu, Rui Liu, Huiyao Inoue, Koji Sumida, Ryuichi Kawahara, Tatsuya Xu, Fan Ye, Lingyu Tian, Wei Kim, Dongjun Jung, Jimin Seo, Jaehyung Wangsajaya, Nadya Yuki Duc, Pham Minh Saxena, Ojasva Nandi, Palash Tao, Xiyan Karlina, Wiwik Luong, Tuan Vasan, Keertana Arun Lee, Roy Ka-Wei Chen, Nancy F. |
| contents | Large language models (LLMs) are now used worldwide, yet their multimodal understanding and reasoning often degrade outside Western, high-resource settings. We propose MMA-ASIA, a comprehensive framework to evaluate LLMs' cultural awareness with a focus on Asian contexts. MMA-ASIA centers on a human-curated, multilingual, and multimodally aligned multiple-choice benchmark covering 8 Asian countries and 10 languages, comprising 27,000 questions; over 79 percent require multi-step reasoning grounded in cultural context, moving beyond simple memorization. To our knowledge, this is the first dataset aligned at the input level across three modalities: text, image (visual question answering), and speech. This enables direct tests of cross-modal transfer. Building on this benchmark, we propose a five-dimensional evaluation protocol that measures: (i) cultural-awareness disparities across countries, (ii) cross-lingual consistency, (iii) cross-modal consistency, (iv) cultural knowledge generalization, and (v) grounding validity. To ensure rigorous assessment, a Cultural Awareness Grounding Validation Module detects "shortcut learning" by checking whether the requisite cultural knowledge supports correct answers. Finally, through comparative model analysis, attention tracing, and an innovative Vision-ablated Prefix Replay (VPR) method, we probe why models diverge across languages and modalities, offering actionable insights for building culturally reliable multimodal LLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_08608 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation Zheng, Weihua Liu, Zhengyuan Chakraborty, Tanmoy Xu, Weiwen Gao, Xiaoxue Tan, Bryan Chen Zhengyu Zou, Bowei Liu, Chang Hu, Yujia Xie, Xing Yi, Xiaoyuan Yao, Jing Wang, Chaojun Li, Long Liu, Rui Liu, Huiyao Inoue, Koji Sumida, Ryuichi Kawahara, Tatsuya Xu, Fan Ye, Lingyu Tian, Wei Kim, Dongjun Jung, Jimin Seo, Jaehyung Wangsajaya, Nadya Yuki Duc, Pham Minh Saxena, Ojasva Nandi, Palash Tao, Xiyan Karlina, Wiwik Luong, Tuan Vasan, Keertana Arun Lee, Roy Ka-Wei Chen, Nancy F. Computation and Language Artificial Intelligence Large language models (LLMs) are now used worldwide, yet their multimodal understanding and reasoning often degrade outside Western, high-resource settings. We propose MMA-ASIA, a comprehensive framework to evaluate LLMs' cultural awareness with a focus on Asian contexts. MMA-ASIA centers on a human-curated, multilingual, and multimodally aligned multiple-choice benchmark covering 8 Asian countries and 10 languages, comprising 27,000 questions; over 79 percent require multi-step reasoning grounded in cultural context, moving beyond simple memorization. To our knowledge, this is the first dataset aligned at the input level across three modalities: text, image (visual question answering), and speech. This enables direct tests of cross-modal transfer. Building on this benchmark, we propose a five-dimensional evaluation protocol that measures: (i) cultural-awareness disparities across countries, (ii) cross-lingual consistency, (iii) cross-modal consistency, (iv) cultural knowledge generalization, and (v) grounding validity. To ensure rigorous assessment, a Cultural Awareness Grounding Validation Module detects "shortcut learning" by checking whether the requisite cultural knowledge supports correct answers. Finally, through comparative model analysis, attention tracing, and an innovative Vision-ablated Prefix Replay (VPR) method, we probe why models diverge across languages and modalities, offering actionable insights for building culturally reliable multimodal LLMs. |
| title | MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2510.08608 |