BenLLMEval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP
Fuente:
arXiv
Saved in:
| Main Authors: | Kabir, Mohsinul, Islam, Mohammed Saidul, Laskar, Md Tahmid Rahman, Nayeem, Mir Tafseer, Bari, M Saiful, Hoque, Enamul |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are Large Vision Language Models up to the Challenge of Chart Comprehension and Reasoning? An Extensive Investigation into the Capabilities and Limitations of LVLMs
by: Islam, Mohammed Saidul, et al.
Published: (2024)
by: Islam, Mohammed Saidul, et al.
Published: (2024)
The Perils of Chart Deception: How Misleading Visualizations Affect Vision-Language Models
by: Mahbub, Ridwan, et al.
Published: (2025)
by: Mahbub, Ridwan, et al.
Published: (2025)
From Charts to Fair Narratives: Uncovering and Mitigating Geo-Economic Biases in Chart-to-Text
by: Mahbub, Ridwan, et al.
Published: (2025)
by: Mahbub, Ridwan, et al.
Published: (2025)
Lost in Translation: Do LVLM Judges Generalize Across Languages?
by: Laskar, Md Tahmid Rahman, et al.
Published: (2026)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2026)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
by: Nayeem, Mir Tafseer, et al.
Published: (2025)
by: Nayeem, Mir Tafseer, et al.
Published: (2025)
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models
by: Alqahtani, Sawsan, et al.
Published: (2026)
by: Alqahtani, Sawsan, et al.
Published: (2026)
Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
Aligning Text, Code, and Vision: A Multi-Objective Reinforcement Learning Framework for Text-to-Visualization
by: Rahman, Mizanur, et al.
Published: (2026)
by: Rahman, Mizanur, et al.
Published: (2026)
DataNarrative: Automated Data-Driven Storytelling with Visualizations and Texts
by: Islam, Mohammed Saidul, et al.
Published: (2024)
by: Islam, Mohammed Saidul, et al.
Published: (2024)
Natural Language Generation for Visualizations: State of the Art, Challenges and Future Directions
by: Hoque, Enamul, et al.
Published: (2024)
by: Hoque, Enamul, et al.
Published: (2024)
LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions
by: Rahman, Mizanur, et al.
Published: (2025)
by: Rahman, Mizanur, et al.
Published: (2025)
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
by: Rahman, Mizanur, et al.
Published: (2025)
by: Rahman, Mizanur, et al.
Published: (2025)
Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
LFOSum: Summarizing Long-form Opinions with Large Language Models
by: Nayeem, Mir Tafseer, et al.
Published: (2024)
by: Nayeem, Mir Tafseer, et al.
Published: (2024)
OpinioRAG: Towards Generating User-Centric Opinion Highlights from Large-scale Online Reviews
by: Nayeem, Mir Tafseer, et al.
Published: (2025)
by: Nayeem, Mir Tafseer, et al.
Published: (2025)
Which English Do LLMs Prefer? Triangulating Structural Bias Towards American English in Foundation Models
by: Nayeem, Mir Tafseer, et al.
Published: (2026)
by: Nayeem, Mir Tafseer, et al.
Published: (2026)
KidLM: Advancing Language Models for Children -- Early Insights and Future Directions
by: Nayeem, Mir Tafseer, et al.
Published: (2024)
by: Nayeem, Mir Tafseer, et al.
Published: (2024)
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations
by: Laskar, Md Tahmid Rahman, et al.
Published: (2024)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2024)
A Comprehensive Evaluation of Large Language Models on Benchmark Biomedical Text Processing Tasks
by: Jahan, Israt, et al.
Published: (2023)
by: Jahan, Israt, et al.
Published: (2023)
Gradient Masters at BLP-2025 Task 1: Advancing Low-Resource NLP for Bengali using Ensemble-Based Adversarial Training for Hate Speech Detection
by: Hoque, Syed Mohaiminul, et al.
Published: (2025)
by: Hoque, Syed Mohaiminul, et al.
Published: (2025)
SurveyGen: Quality-Aware Scientific Survey Generation with Large Language Models
by: Bao, Tong, et al.
Published: (2025)
by: Bao, Tong, et al.
Published: (2025)
Islamic Lifestyle Applications: Meeting the Spiritual Needs of Modern Muslims
by: Kabir, Mohsinul, et al.
Published: (2024)
by: Kabir, Mohsinul, et al.
Published: (2024)
DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
by: Kartha, Aaryaman, et al.
Published: (2025)
by: Kartha, Aaryaman, et al.
Published: (2025)
Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
by: Jahan, Israt, et al.
Published: (2025)
by: Jahan, Israt, et al.
Published: (2025)
XL-HeadTags: Leveraging Multimodal Retrieval Augmentation for the Multilingual Generation of News Headlines and Tags
by: Shohan, Faisal Tareque, et al.
Published: (2024)
by: Shohan, Faisal Tareque, et al.
Published: (2024)
SparseTransX: Efficient Training of Translation-Based Knowledge Graph Embeddings Using Sparse Matrix Operations
by: Anik, Md Saidul Hoque, et al.
Published: (2025)
by: Anik, Md Saidul Hoque, et al.
Published: (2025)
State and politics in the transitional era of globalization: Twisting and turning toward authoritarian and hybrid regimes
by: Hafijur Rahman, et al.
Published: (2024)
by: Hafijur Rahman, et al.
Published: (2024)
Bengali Text Classification: An Evaluation of Large Language Model Approaches
by: Hoque, Md Mahmudul, et al.
Published: (2026)
by: Hoque, Md Mahmudul, et al.
Published: (2026)
A Comprehensive Survey on Visual Question Answering Datasets and Algorithms
by: Kabir, Raihan, et al.
Published: (2024)
by: Kabir, Raihan, et al.
Published: (2024)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
by: Kabir, Mohsinul, et al.
Published: (2025)
by: Kabir, Mohsinul, et al.
Published: (2025)
PULSAR: Graph based Positive Unlabeled Learning with Multi Stream Adaptive Convolutions for Parkinson's Disease Recognition
by: Alam, Md. Zarif Ul, et al.
Published: (2023)
by: Alam, Md. Zarif Ul, et al.
Published: (2023)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
by: Kabir, Mohsinul, et al.
Published: (2026)
by: Kabir, Mohsinul, et al.
Published: (2026)
eC-Tab2Text: Aspect-Based Text Generation from e-Commerce Product Tables
by: Guanilo, Luis Antonio Gutiérrez, et al.
Published: (2025)
by: Guanilo, Luis Antonio Gutiérrez, et al.
Published: (2025)
Managing risk and reaping rewards: Climate‐change futures as a game‐changer for energy futures markets
by: Mohammad Enamul Hoque, et al.
Published: (2024)
by: Mohammad Enamul Hoque, et al.
Published: (2024)
Semantic Label Drift in Cross-Cultural Translation
by: Kabir, Mohsinul, et al.
Published: (2025)
by: Kabir, Mohsinul, et al.
Published: (2025)
S$^3$F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network
by: Siddiqui, Md. Saiful Bari, et al.
Published: (2025)
by: Siddiqui, Md. Saiful Bari, et al.
Published: (2025)
Two Decades of Bengali Handwritten Digit Recognition: A Survey
by: Rahman, A. B. M. Ashikur, et al.
Published: (2022)
by: Rahman, A. B. M. Ashikur, et al.
Published: (2022)
BanglaIPA: Towards Robust Text-to-IPA Transcription with Contextual Rewriting in Bengali
by: Hasan, Jakir, et al.
Published: (2026)
by: Hasan, Jakir, et al.
Published: (2026)
Similar Items
-
Are Large Vision Language Models up to the Challenge of Chart Comprehension and Reasoning? An Extensive Investigation into the Capabilities and Limitations of LVLMs
by: Islam, Mohammed Saidul, et al.
Published: (2024) -
The Perils of Chart Deception: How Misleading Visualizations Affect Vision-Language Models
by: Mahbub, Ridwan, et al.
Published: (2025) -
From Charts to Fair Narratives: Uncovering and Mitigating Geo-Economic Biases in Chart-to-Text
by: Mahbub, Ridwan, et al.
Published: (2025) -
Lost in Translation: Do LVLM Judges Generalize Across Languages?
by: Laskar, Md Tahmid Rahman, et al.
Published: (2026) -
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)