Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Bhowmik, Shimanto, Dipto, Tawsif Tashwar, Islam, Md Sazzad, Hsu, Sheryl, Reasat, Tahsin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The "Colonial Impulse" of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based Biases
by: Das, Dipto, et al.
Published: (2024)
by: Das, Dipto, et al.
Published: (2024)
Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages?
by: Dipto, Tawsif Tashwar, et al.
Published: (2025)
by: Dipto, Tawsif Tashwar, et al.
Published: (2025)
RegSpeech12: A Regional Corpus of Bengali Spontaneous Speech Across Dialects
by: Hassan, Md. Rezuwan, et al.
Published: (2025)
by: Hassan, Md. Rezuwan, et al.
Published: (2025)
Parameter-Efficient Fine-Tuning for Low-Resource Languages: A Comparative Study of LLMs for Bengali Hate Speech Detection
by: Islam, Akif, et al.
Published: (2025)
by: Islam, Akif, et al.
Published: (2025)
From Lightweight CNNs to SpikeNets: Benchmarking Accuracy-Energy Tradeoffs with Pruned Spiking SqueezeNet
by: Kabir, Radib Bin, et al.
Published: (2026)
by: Kabir, Radib Bin, et al.
Published: (2026)
FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models
by: Mateega, Spencer, et al.
Published: (2025)
by: Mateega, Spencer, et al.
Published: (2025)
BTPD: A Multilingual Hand-curated Dataset of Bengali Transnational Political Discourse Across Online Communities
by: Das, Dipto, et al.
Published: (2025)
by: Das, Dipto, et al.
Published: (2025)
Risks, Causes, and Mitigations of Widespread Deployments of Large Language Models (LLMs): A Survey
by: Sakib, Md Nazmus, et al.
Published: (2024)
by: Sakib, Md Nazmus, et al.
Published: (2024)
MultiSoc-4D: A Benchmark for Diagnosing Instruction-Induced Label Collapse in Closed-Set LLM Annotation of Bengali Social Media
by: Pramanik, Souvik, et al.
Published: (2026)
by: Pramanik, Souvik, et al.
Published: (2026)
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
by: Li, Haoming, et al.
Published: (2025)
by: Li, Haoming, et al.
Published: (2025)
From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge
by: Chowdhury, Nafis, et al.
Published: (2025)
by: Chowdhury, Nafis, et al.
Published: (2025)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026)
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026)
BeliN: A Novel Corpus for Bengali Religious News Headline Generation using Contextual Feature Fusion
by: Osama, Md, et al.
Published: (2025)
by: Osama, Md, et al.
Published: (2025)
Performance Analysis of Few-Shot Learning Approaches for Bangla Handwritten Character and Digit Recognition
by: Ahamed, Mehedi, et al.
Published: (2025)
by: Ahamed, Mehedi, et al.
Published: (2025)
RTP-LX: Can LLMs Evaluate Toxicity in Multilingual Scenarios?
by: de Wynter, Adrian, et al.
Published: (2024)
by: de Wynter, Adrian, et al.
Published: (2024)
Sentiment Polarity Analysis of Bangla Food Reviews Using Machine and Deep Learning Algorithms
by: Amin, Al, et al.
Published: (2024)
by: Amin, Al, et al.
Published: (2024)
Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models
by: Chua, Lynn, et al.
Published: (2024)
by: Chua, Lynn, et al.
Published: (2024)
An Image Dataset of Common Skin Diseases of Bangladesh and Benchmarking Performance with Machine Learning Models
by: Hossain, Sazzad, et al.
Published: (2026)
by: Hossain, Sazzad, et al.
Published: (2026)
Automatic Pull Request Description Generation Using LLMs: A T5 Model Approach
by: Sakib, Md Nazmus, et al.
Published: (2024)
by: Sakib, Md Nazmus, et al.
Published: (2024)
Ensemble Language Models for Multilingual Sentiment Analysis
by: Hasan, Md Arid
Published: (2024)
by: Hasan, Md Arid
Published: (2024)
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
by: Liu, Yixin, et al.
Published: (2023)
by: Liu, Yixin, et al.
Published: (2023)
On the Calibration of Multilingual Question Answering LLMs
by: Yang, Yahan, et al.
Published: (2023)
by: Yang, Yahan, et al.
Published: (2023)
Evaluating LLMs' Reasoning Over Ordered Procedural Steps
by: Anika, Adrita, et al.
Published: (2025)
by: Anika, Adrita, et al.
Published: (2025)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
by: Liu, Yijun, et al.
Published: (2024)
by: Liu, Yijun, et al.
Published: (2024)
Better To Ask in English? Evaluating Factual Accuracy of Multilingual LLMs in English and Low-Resource Languages
by: Rohera, Pritika, et al.
Published: (2025)
by: Rohera, Pritika, et al.
Published: (2025)
Explainable Multimodal Sentiment Analysis on Bengali Memes
by: Elahi, Kazi Toufique, et al.
Published: (2023)
by: Elahi, Kazi Toufique, et al.
Published: (2023)
LLM-Mixer: Multiscale Mixing in LLMs for Time Series Forecasting
by: Kowsher, Md, et al.
Published: (2024)
by: Kowsher, Md, et al.
Published: (2024)
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
by: Nyandwi, Jean de Dieu, et al.
Published: (2025)
by: Nyandwi, Jean de Dieu, et al.
Published: (2025)
Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
How Does Quantization Affect Multilingual LLMs?
by: Marchisio, Kelly, et al.
Published: (2024)
by: Marchisio, Kelly, et al.
Published: (2024)
Empowering Bengali Education with AI: Solving Bengali Math Word Problems through Transformer Models
by: Era, Jalisha Jashim, et al.
Published: (2025)
by: Era, Jalisha Jashim, et al.
Published: (2025)
Bridging the Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
by: Kumar, Somnath, et al.
Published: (2024)
by: Kumar, Somnath, et al.
Published: (2024)
Modeling Romanized Hindi and Bengali: Dataset Creation and Multilingual LLM Integration
by: Gharami, Kanchon, et al.
Published: (2025)
by: Gharami, Kanchon, et al.
Published: (2025)
Dhoroni: Exploring Bengali Climate Change and Environmental Views with a Multi-Perspective News Dataset and Natural Language Processing
by: Wasi, Azmine Toushik, et al.
Published: (2024)
by: Wasi, Azmine Toushik, et al.
Published: (2024)
BanglaSentNet: An Explainable Hybrid Deep Learning Framework for Multi-Aspect Sentiment Analysis with Cross-Domain Transfer Learning
by: Islam, Ariful, et al.
Published: (2025)
by: Islam, Ariful, et al.
Published: (2025)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
by: Oh, Jungwoo, et al.
Published: (2026)
by: Oh, Jungwoo, et al.
Published: (2026)
MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation
by: Agarwal, Mehul, et al.
Published: (2026)
by: Agarwal, Mehul, et al.
Published: (2026)
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
by: Longpre, Shayne, et al.
Published: (2025)
by: Longpre, Shayne, et al.
Published: (2025)
AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Margin
by: Xiong, Jian, et al.
Published: (2025)
by: Xiong, Jian, et al.
Published: (2025)
Iterative Self-Tuning LLMs for Enhanced Jailbreaking Capabilities
by: Sun, Chung-En, et al.
Published: (2024)
by: Sun, Chung-En, et al.
Published: (2024)
Similar Items
-
The "Colonial Impulse" of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based Biases
by: Das, Dipto, et al.
Published: (2024) -
Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages?
by: Dipto, Tawsif Tashwar, et al.
Published: (2025) -
RegSpeech12: A Regional Corpus of Bengali Spontaneous Speech Across Dialects
by: Hassan, Md. Rezuwan, et al.
Published: (2025) -
Parameter-Efficient Fine-Tuning for Low-Resource Languages: A Comparative Study of LLMs for Bengali Hate Speech Detection
by: Islam, Akif, et al.
Published: (2025) -
From Lightweight CNNs to SpikeNets: Benchmarking Accuracy-Energy Tradeoffs with Pruned Spiking SqueezeNet
by: Kabir, Radib Bin, et al.
Published: (2026)