Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhowmik, Shimanto, Dipto, Tawsif Tashwar, Islam, Md Sazzad, Hsu, Sheryl, Reasat, Tahsin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908473490407424
author Bhowmik, Shimanto
Dipto, Tawsif Tashwar
Islam, Md Sazzad
Hsu, Sheryl
Reasat, Tahsin
author_facet Bhowmik, Shimanto
Dipto, Tawsif Tashwar
Islam, Md Sazzad
Hsu, Sheryl
Reasat, Tahsin
contents Bengali is an underrepresented language in NLP research. However, it remains a challenge due to its unique linguistic structure and computational constraints. In this work, we systematically investigate the challenges that hinder Bengali NLP performance by focusing on the absence of standardized evaluation benchmarks. We then evaluated 10 recent open source Large Language Models (LLMs) in 8 of the translated datasets and performed a comprehensive error analysis to pinpoint their primary failure modes. Our findings reveal consistent performance gaps for Bengali compared to English, particularly for smaller models and specific model families like Mistral. We also identified promising robustness in certain architectures, such as DeepSeek, that maintain more stable performance across languages. Our analysis reveals an inverse relationship between tokenization efficiency and LLM accuracy where models tend to perform worse when inputs are excessively tokenized, whereas more efficient \& concise tokenization results in improved performance. These findings highlight critical areas where current models fall short and underscore the need for improved dataset quality and evaluation methodologies tailored to multilingual contexts. This work will catalyze further research on NLP for underrepresented languages, helping to democratize access to advanced language technologies worldwide. The code and dataset used in this research is publicly available at https://github.com/BengaliAI/bn-llm-benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23248
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis
Bhowmik, Shimanto
Dipto, Tawsif Tashwar
Islam, Md Sazzad
Hsu, Sheryl
Reasat, Tahsin
Computation and Language
Machine Learning
Bengali is an underrepresented language in NLP research. However, it remains a challenge due to its unique linguistic structure and computational constraints. In this work, we systematically investigate the challenges that hinder Bengali NLP performance by focusing on the absence of standardized evaluation benchmarks. We then evaluated 10 recent open source Large Language Models (LLMs) in 8 of the translated datasets and performed a comprehensive error analysis to pinpoint their primary failure modes. Our findings reveal consistent performance gaps for Bengali compared to English, particularly for smaller models and specific model families like Mistral. We also identified promising robustness in certain architectures, such as DeepSeek, that maintain more stable performance across languages. Our analysis reveals an inverse relationship between tokenization efficiency and LLM accuracy where models tend to perform worse when inputs are excessively tokenized, whereas more efficient \& concise tokenization results in improved performance. These findings highlight critical areas where current models fall short and underscore the need for improved dataset quality and evaluation methodologies tailored to multilingual contexts. This work will catalyze further research on NLP for underrepresented languages, helping to democratize access to advanced language technologies worldwide. The code and dataset used in this research is publicly available at https://github.com/BengaliAI/bn-llm-benchmark.
title Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2507.23248