mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Beyene, Luel Hagos, Verma, Vivek, Ma, Min, Alabi, Jesujoba O., Schmidt, Fabian David, Nakatumba-Nabende, Joyce, Adelani, David Ifeoluwa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914009339396096
author Beyene, Luel Hagos
Verma, Vivek
Ma, Min
Alabi, Jesujoba O.
Schmidt, Fabian David
Nakatumba-Nabende, Joyce
Adelani, David Ifeoluwa
author_facet Beyene, Luel Hagos
Verma, Vivek
Ma, Min
Alabi, Jesujoba O.
Schmidt, Fabian David
Nakatumba-Nabende, Joyce
Adelani, David Ifeoluwa
contents Large Language models (LLMs) have demonstrated impressive performance on a wide range of tasks, including in multimodal settings such as speech. However, their evaluation is often limited to English and a few high-resource languages. For low-resource languages, there is no standardized evaluation benchmark. In this paper, we address this gap by introducing mSTEB, a new benchmark to evaluate the performance of LLMs on a wide range of tasks covering language identification, text classification, question answering, and translation tasks on both speech and text modalities. We evaluated the performance of leading LLMs such as Gemini 2.0 Flash and GPT-4o (Audio) and state-of-the-art open models such as Qwen 2 Audio and Gemma 3 27B. Our evaluation shows a wide gap in performance between high-resource and low-resource languages, especially for languages spoken in Africa and Americas/Oceania. Our findings show that more investment is needed to address their under-representation in LLMs coverage.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08400
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
Beyene, Luel Hagos
Verma, Vivek
Ma, Min
Alabi, Jesujoba O.
Schmidt, Fabian David
Nakatumba-Nabende, Joyce
Adelani, David Ifeoluwa
Computation and Language
Machine Learning
Sound
Audio and Speech Processing
Large Language models (LLMs) have demonstrated impressive performance on a wide range of tasks, including in multimodal settings such as speech. However, their evaluation is often limited to English and a few high-resource languages. For low-resource languages, there is no standardized evaluation benchmark. In this paper, we address this gap by introducing mSTEB, a new benchmark to evaluate the performance of LLMs on a wide range of tasks covering language identification, text classification, question answering, and translation tasks on both speech and text modalities. We evaluated the performance of leading LLMs such as Gemini 2.0 Flash and GPT-4o (Audio) and state-of-the-art open models such as Qwen 2 Audio and Gemma 3 27B. Our evaluation shows a wide gap in performance between high-resource and low-resource languages, especially for languages spoken in Africa and Americas/Oceania. Our findings show that more investment is needed to address their under-representation in LLMs coverage.
title mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
topic Computation and Language
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.08400