Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yao, Jui-Ming, Xie, Bing-Cheng, Peng, Sheng-Wei, Chen, Hao-Yuan, Zheng, He-Rong, Tan, Bing-Jia, Wang, Peter Shaojui, Su, Shun-Feng
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913972189396992
author Yao, Jui-Ming
Xie, Bing-Cheng
Peng, Sheng-Wei
Chen, Hao-Yuan
Zheng, He-Rong
Tan, Bing-Jia
Wang, Peter Shaojui
Su, Shun-Feng
author_facet Yao, Jui-Ming
Xie, Bing-Cheng
Peng, Sheng-Wei
Chen, Hao-Yuan
Zheng, He-Rong
Tan, Bing-Jia
Wang, Peter Shaojui
Su, Shun-Feng
contents Multimodal Large Language Models (MLLMs) process visual, acoustic, and textual inputs, addressing the limitations of single-modality LLMs. However, existing benchmarks often overlook tri-modal evaluation in Traditional Chinese and do not consider inference latency. To address this, we introduce Multi-TW, the first Traditional Chinese benchmark for evaluating the performance and latency of any-to-any multimodal models. Multi-TW includes 900 multiple-choice questions (image and text, audio and text pairs) sourced from official proficiency tests developed with the Steering Committee for the Test of Proficiency-Huayu (SC-TOP). We evaluated various any-to-any models and vision-language models (VLMs) with audio transcription. Our results show that closed-source models generally outperform open-source ones across modalities, although open-source models can perform well in audio tasks. End-to-end any-to-any pipelines offer clear latency advantages compared to VLMs using separate audio transcription. Multi-TW presents a comprehensive view of model capabilities and highlights the need for Traditional Chinese fine-tuning and efficient multimodal architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01274
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan
Yao, Jui-Ming
Xie, Bing-Cheng
Peng, Sheng-Wei
Chen, Hao-Yuan
Zheng, He-Rong
Tan, Bing-Jia
Wang, Peter Shaojui
Su, Shun-Feng
Artificial Intelligence
Computation and Language
Multimodal Large Language Models (MLLMs) process visual, acoustic, and textual inputs, addressing the limitations of single-modality LLMs. However, existing benchmarks often overlook tri-modal evaluation in Traditional Chinese and do not consider inference latency. To address this, we introduce Multi-TW, the first Traditional Chinese benchmark for evaluating the performance and latency of any-to-any multimodal models. Multi-TW includes 900 multiple-choice questions (image and text, audio and text pairs) sourced from official proficiency tests developed with the Steering Committee for the Test of Proficiency-Huayu (SC-TOP). We evaluated various any-to-any models and vision-language models (VLMs) with audio transcription. Our results show that closed-source models generally outperform open-source ones across modalities, although open-source models can perform well in audio tasks. End-to-end any-to-any pipelines offer clear latency advantages compared to VLMs using separate audio transcription. Multi-TW presents a comprehensive view of model capabilities and highlights the need for Traditional Chinese fine-tuning and efficient multimodal architectures.
title Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.01274