A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations
Fuente:
arXiv
Saved in:
| Main Authors: | Laskar, Md Tahmid Rahman, Alqahtani, Sawsan, Bari, M Saiful, Rahman, Mizanur, Khan, Mohammad Abdullah Matin, Khan, Haidar, Jahan, Israt, Bhuiyan, Amran, Tan, Chee Wei, Parvez, Md Rizwan, Hoque, Enamul, Joty, Shafiq, Huang, Jimmy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
by: Rahman, Mizanur, et al.
Published: (2025)
by: Rahman, Mizanur, et al.
Published: (2025)
Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
DataNarrative: Automated Data-Driven Storytelling with Visualizations and Texts
by: Islam, Mohammed Saidul, et al.
Published: (2024)
by: Islam, Mohammed Saidul, et al.
Published: (2024)
Lost in Translation: Do LVLM Judges Generalize Across Languages?
by: Laskar, Md Tahmid Rahman, et al.
Published: (2026)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2026)
Aligning Text, Code, and Vision: A Multi-Objective Reinforcement Learning Framework for Text-to-Visualization
by: Rahman, Mizanur, et al.
Published: (2026)
by: Rahman, Mizanur, et al.
Published: (2026)
LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions
by: Rahman, Mizanur, et al.
Published: (2025)
by: Rahman, Mizanur, et al.
Published: (2025)
Evolution of ReID: From Early Methods to LLM Integration
by: Bhuiyan, Amran, et al.
Published: (2025)
by: Bhuiyan, Amran, et al.
Published: (2025)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
From Charts to Fair Narratives: Uncovering and Mitigating Geo-Economic Biases in Chart-to-Text
by: Mahbub, Ridwan, et al.
Published: (2025)
by: Mahbub, Ridwan, et al.
Published: (2025)
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models
by: Alqahtani, Sawsan, et al.
Published: (2026)
by: Alqahtani, Sawsan, et al.
Published: (2026)
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
by: Nayeem, Mir Tafseer, et al.
Published: (2025)
by: Nayeem, Mir Tafseer, et al.
Published: (2025)
A Comprehensive Evaluation of Large Language Models on Benchmark Biomedical Text Processing Tasks
by: Jahan, Israt, et al.
Published: (2023)
by: Jahan, Israt, et al.
Published: (2023)
Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
by: Jahan, Israt, et al.
Published: (2025)
by: Jahan, Israt, et al.
Published: (2025)
ChartInstruct: Instruction Tuning for Chart Comprehension and Reasoning
by: Masry, Ahmed, et al.
Published: (2024)
by: Masry, Ahmed, et al.
Published: (2024)
ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
Open-RAG: Enhanced Retrieval-Augmented Reasoning with Open-Source Large Language Models
by: Islam, Shayekh Bin, et al.
Published: (2024)
by: Islam, Shayekh Bin, et al.
Published: (2024)
The Perils of Chart Deception: How Misleading Visualizations Affect Vision-Language Models
by: Mahbub, Ridwan, et al.
Published: (2025)
by: Mahbub, Ridwan, et al.
Published: (2025)
BenLLMEval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP
by: Kabir, Mohsinul, et al.
Published: (2023)
by: Kabir, Mohsinul, et al.
Published: (2023)
DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
by: Kartha, Aaryaman, et al.
Published: (2025)
by: Kartha, Aaryaman, et al.
Published: (2025)
DATAREEL: Automated Data-Driven Video Story Generation with Animations
by: Mahbub, Ridwan, et al.
Published: (2026)
by: Mahbub, Ridwan, et al.
Published: (2026)
Are Large Vision Language Models up to the Challenge of Chart Comprehension and Reasoning? An Extensive Investigation into the Capabilities and Limitations of LVLMs
by: Islam, Mohammed Saidul, et al.
Published: (2024)
by: Islam, Mohammed Saidul, et al.
Published: (2024)
Utilizing BERT for Information Retrieval: Survey, Applications, Resources, and Challenges
by: Wang, Jiajia, et al.
Published: (2024)
by: Wang, Jiajia, et al.
Published: (2024)
3D‐structure coupled monopole designs with lower frequency resonance
by: Enamul Khan, et al.
Published: (2024)
by: Enamul Khan, et al.
Published: (2024)
Impact of Salinisation on the Neighbour-based Spatial Diversity of Tree Species in the Sundarbans Mangrove of Bangladesh
by: Rahman, Md Mizanur
Published: (2026)
by: Rahman, Md Mizanur
Published: (2026)
An assessment of agroforestry as a climate‐smart practice: Evidences from farmers of northwestern region of Bangladesh
by: Md. Manik Ali, et al.
Published: (2024)
by: Md. Manik Ali, et al.
Published: (2024)
Xolver: Multi-Agent Reasoning with Holistic Experience Learning Just Like an Olympiad Team
by: Hosain, Md Tanzib, et al.
Published: (2025)
by: Hosain, Md Tanzib, et al.
Published: (2025)
S$^3$F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network
by: Siddiqui, Md. Saiful Bari, et al.
Published: (2025)
by: Siddiqui, Md. Saiful Bari, et al.
Published: (2025)
Chain of Evidences and Evidence to Generate: Prompting for Context Grounded and Retrieval Augmented Reasoning
by: Parvez, Md Rizwan
Published: (2024)
by: Parvez, Md Rizwan
Published: (2024)
Position: Beyond Assistance -- Reimagining LLMs as Ethical and Adaptive Co-Creators in Mental Health Care
by: Badawi, Abeer, et al.
Published: (2025)
by: Badawi, Abeer, et al.
Published: (2025)
Madhhab-Based Differences in Fiqh Verses of the Qur'an: A Comparative Analysis of Hanafi and Hanbali Perspectives in Shariah Interpretation
by: Khan, Md Taki Tahmid Khan
Published: (2025)
by: Khan, Md Taki Tahmid Khan
Published: (2025)
A Heterogeneous Two-Stream Framework for Video Action Recognition with Comparative Fusion Analysis
by: Rahaman, Md. Afzalur, et al.
Published: (2026)
by: Rahaman, Md. Afzalur, et al.
Published: (2026)
Reward Engineering for Reinforcement Learning in Software Tasks
by: Masud, Md Rayhanul, et al.
Published: (2026)
by: Masud, Md Rayhanul, et al.
Published: (2026)
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
by: Paul, Dhiman, et al.
Published: (2024)
by: Paul, Dhiman, et al.
Published: (2024)
Melamine‐Modified Waterborne Polyurethane Coatings: Experimental and DFT Study
by: Urbana Kawsar Mitali, et al.
Published: (2025)
by: Urbana Kawsar Mitali, et al.
Published: (2025)
Comparative Evaluation of the Zootechnical Performance, Body Composition, Haemato‐Biochemical Profile, and Enzymatic Activity of Nile tilapia ( Oreochromis niloticus ) Cultured in Biofloc and Conventional Rearing Systems
by: Sharmin Aktar, et al.
Published: (2025)
by: Sharmin Aktar, et al.
Published: (2025)
A Survey on Agentic Security: Applications, Threats and Defenses
by: Shahriar, Asif, et al.
Published: (2025)
by: Shahriar, Asif, et al.
Published: (2025)
Future Mining: Learning for Safety and Security
by: Rahman, Md Sazedur, et al.
Published: (2026)
by: Rahman, Md Sazedur, et al.
Published: (2026)
PULSAR: Graph based Positive Unlabeled Learning with Multi Stream Adaptive Convolutions for Parkinson's Disease Recognition
by: Alam, Md. Zarif Ul, et al.
Published: (2023)
by: Alam, Md. Zarif Ul, et al.
Published: (2023)
Security Vulnerabilities in Software Supply Chain for Autonomous Vehicles
by: Haque, Md Wasiul, et al.
Published: (2025)
by: Haque, Md Wasiul, et al.
Published: (2025)
Similar Items
-
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
by: Rahman, Mizanur, et al.
Published: (2025) -
Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025) -
DataNarrative: Automated Data-Driven Storytelling with Visualizations and Texts
by: Islam, Mohammed Saidul, et al.
Published: (2024) -
Lost in Translation: Do LVLM Judges Generalize Across Languages?
by: Laskar, Md Tahmid Rahman, et al.
Published: (2026) -
Aligning Text, Code, and Vision: A Multi-Objective Reinforcement Learning Framework for Text-to-Visualization
by: Rahman, Mizanur, et al.
Published: (2026)