Where Are We? Evaluating LLM Performance on African Languages
Fuente:
arXiv
Saved in:
| Main Authors: | Adebara, Ife, Toyin, Hawau Olamide, Ghebremichael, Nahom Tesfu, Elmadany, AbdelRahim, Abdul-Mageed, Muhammad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cheetah: Natural Language Generation for 517 African Languages
by: Adebara, Ife, et al.
Published: (2024)
by: Adebara, Ife, et al.
Published: (2024)
Toucan: Many-to-Many Translation for 150 African Language Pairs
by: Elmadany, AbdelRahim, et al.
Published: (2024)
by: Elmadany, AbdelRahim, et al.
Published: (2024)
Voice of a Continent: Mapping Africa's Speech Technology Frontier
by: Elmadany, AbdelRahim, et al.
Published: (2025)
by: Elmadany, AbdelRahim, et al.
Published: (2025)
AfroScope: A Framework for Studying the Linguistic Landscape of Africa
by: Kwon, Sang Yun, et al.
Published: (2026)
by: Kwon, Sang Yun, et al.
Published: (2026)
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology
by: Sullivan, Peter, et al.
Published: (2026)
by: Sullivan, Peter, et al.
Published: (2026)
NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task
by: Talafha, Bashar, et al.
Published: (2025)
by: Talafha, Bashar, et al.
Published: (2025)
Interplay of Machine Translation, Diacritics, and Diacritization
by: Chen, Wei-Rui, et al.
Published: (2024)
by: Chen, Wei-Rui, et al.
Published: (2024)
WojoodNER 2024: The Second Arabic Named Entity Recognition Shared Task
by: Jarrar, Mustafa, et al.
Published: (2024)
by: Jarrar, Mustafa, et al.
Published: (2024)
Fumbling in Babel: An Investigation into ChatGPT's Language Identification Ability
by: Chen, Wei-Rui, et al.
Published: (2023)
by: Chen, Wei-Rui, et al.
Published: (2023)
NADI 2024: The Fifth Nuanced Arabic Dialect Identification Shared Task
by: Abdul-Mageed, Muhammad, et al.
Published: (2024)
by: Abdul-Mageed, Muhammad, et al.
Published: (2024)
Are LLMs Good Text Diacritizers? An Arabic and Yoruba Case Study
by: Toyin, Hawau Olamide, et al.
Published: (2025)
by: Toyin, Hawau Olamide, et al.
Published: (2025)
STTATTS: Unified Speech-To-Text And Text-To-Speech Model
by: Toyin, Hawau Olamide, et al.
Published: (2024)
by: Toyin, Hawau Olamide, et al.
Published: (2024)
Dialectal Coverage And Generalization in Arabic Speech Recognition
by: Djanibekov, Amirbek, et al.
Published: (2024)
by: Djanibekov, Amirbek, et al.
Published: (2024)
Exploring the Limitations of Detecting Machine-Generated Text
by: Doughman, Jad, et al.
Published: (2024)
by: Doughman, Jad, et al.
Published: (2024)
ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
by: Toyin, Hawau Olamide, et al.
Published: (2025)
by: Toyin, Hawau Olamide, et al.
Published: (2025)
LLM Performance Predictors are good initializers for Architecture Search
by: Jawahar, Ganesh, et al.
Published: (2023)
by: Jawahar, Ganesh, et al.
Published: (2023)
Aligning Stuttered-Speech Research with End-User Needs: Scoping Review, Survey, and Guidelines
by: Toyin, Hawau Olamide, et al.
Published: (2026)
by: Toyin, Hawau Olamide, et al.
Published: (2026)
Dallah: A Dialect-Aware Multimodal Large Language Model for Arabic
by: Alwajih, Fakhraddin, et al.
Published: (2024)
by: Alwajih, Fakhraddin, et al.
Published: (2024)
Effective Self-Mining of In-Context Examples for Unsupervised Machine Translation with LLMs
by: Mekki, Abdellah El, et al.
Published: (2024)
by: Mekki, Abdellah El, et al.
Published: (2024)
Clinical Annotations for Automatic Stuttering Severity Assessment
by: Valente, Ana Rita, et al.
Published: (2025)
by: Valente, Ana Rita, et al.
Published: (2025)
DetoxLLM: A Framework for Detoxification with Explanations
by: Khondaker, Md Tawkat Islam, et al.
Published: (2024)
by: Khondaker, Md Tawkat Islam, et al.
Published: (2024)
Autoregressive + Chain of Thought = Recurrent: Recurrence's Role in Language Models' Computability and a Revisit of Recurrent Transformer
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
To Distill or Not to Distill? On the Robustness of Robust Knowledge Distillation
by: Waheed, Abdul, et al.
Published: (2024)
by: Waheed, Abdul, et al.
Published: (2024)
Arabic Automatic Story Generation with Large Language Models
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
On Barriers to Archival Audio Processing
by: Sullivan, Peter, et al.
Published: (2025)
by: Sullivan, Peter, et al.
Published: (2025)
Gazelle: An Instruction Dataset for Arabic Writing Assistance
by: Magdy, Samar M., et al.
Published: (2024)
by: Magdy, Samar M., et al.
Published: (2024)
Zero-Shot Context-Aware ASR for Diverse Arabic Varieties
by: Talafha, Bashar, et al.
Published: (2025)
by: Talafha, Bashar, et al.
Published: (2025)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
by: Doan, Khai Duy, et al.
Published: (2024)
by: Doan, Khai Duy, et al.
Published: (2024)
Qalam : A Multimodal LLM for Arabic Optical Character and Handwriting Recognition
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes
by: Waheed, Abdul, et al.
Published: (2024)
by: Waheed, Abdul, et al.
Published: (2024)
LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions
by: Wu, Minghao, et al.
Published: (2023)
by: Wu, Minghao, et al.
Published: (2023)
FinTral: A Family of GPT-4 Level Multimodal Financial Large Language Models
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
Where Are We At with Automatic Speech Recognition for the Bambara Language?
by: Diallo, Seydou, et al.
Published: (2026)
by: Diallo, Seydou, et al.
Published: (2026)
DefenderBench: A Toolkit for Evaluating Language Agents in Cybersecurity Environments
by: Zhang, Chiyu, et al.
Published: (2025)
by: Zhang, Chiyu, et al.
Published: (2025)
Peacock: A Family of Arabic Multimodal Large Language Models and Benchmarks
by: Alwajih, Fakhraddin, et al.
Published: (2024)
by: Alwajih, Fakhraddin, et al.
Published: (2024)
Jawaher: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking
by: Magdy, Samar M., et al.
Published: (2025)
by: Magdy, Samar M., et al.
Published: (2025)
NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities
by: Mekki, Abdellah El, et al.
Published: (2025)
by: Mekki, Abdellah El, et al.
Published: (2025)
Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages
by: Almheiri, Saeed, et al.
Published: (2026)
by: Almheiri, Saeed, et al.
Published: (2026)
Beyond Content: How Grammatical Gender Shapes Visual Representation in Text-to-Image Models
by: Saeed, Muhammed, et al.
Published: (2025)
by: Saeed, Muhammed, et al.
Published: (2025)
Similar Items
-
Cheetah: Natural Language Generation for 517 African Languages
by: Adebara, Ife, et al.
Published: (2024) -
Toucan: Many-to-Many Translation for 150 African Language Pairs
by: Elmadany, AbdelRahim, et al.
Published: (2024) -
Voice of a Continent: Mapping Africa's Speech Technology Frontier
by: Elmadany, AbdelRahim, et al.
Published: (2025) -
AfroScope: A Framework for Studying the Linguistic Landscape of Africa
by: Kwon, Sang Yun, et al.
Published: (2026) -
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology
by: Sullivan, Peter, et al.
Published: (2026)