Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
Fuente:
arXiv
Salvato in:
| Autori principali: | Laskar, Md Tahmid Rahman, Islam, Mohammed Saidul, Mahbub, Ridwan, Rahman, Mizanur, Bhuiyan, Amran, Jahan, Israt, Nayeem, Mir Tafseer, Joty, Shafiq, Hoque, Enamul, Huang, Jimmy |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Lost in Translation: Do LVLM Judges Generalize Across Languages?
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2026)
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2026)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
From Charts to Fair Narratives: Uncovering and Mitigating Geo-Economic Biases in Chart-to-Text
di: Mahbub, Ridwan, et al.
Pubblicazione: (2025)
di: Mahbub, Ridwan, et al.
Pubblicazione: (2025)
The Perils of Chart Deception: How Misleading Visualizations Affect Vision-Language Models
di: Mahbub, Ridwan, et al.
Pubblicazione: (2025)
di: Mahbub, Ridwan, et al.
Pubblicazione: (2025)
LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions
di: Rahman, Mizanur, et al.
Pubblicazione: (2025)
di: Rahman, Mizanur, et al.
Pubblicazione: (2025)
Aligning Text, Code, and Vision: A Multi-Objective Reinforcement Learning Framework for Text-to-Visualization
di: Rahman, Mizanur, et al.
Pubblicazione: (2026)
di: Rahman, Mizanur, et al.
Pubblicazione: (2026)
Are Large Vision Language Models up to the Challenge of Chart Comprehension and Reasoning? An Extensive Investigation into the Capabilities and Limitations of LVLMs
di: Islam, Mohammed Saidul, et al.
Pubblicazione: (2024)
di: Islam, Mohammed Saidul, et al.
Pubblicazione: (2024)
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
di: Rahman, Mizanur, et al.
Pubblicazione: (2025)
di: Rahman, Mizanur, et al.
Pubblicazione: (2025)
DATAREEL: Automated Data-Driven Video Story Generation with Animations
di: Mahbub, Ridwan, et al.
Pubblicazione: (2026)
di: Mahbub, Ridwan, et al.
Pubblicazione: (2026)
BenLLMEval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP
di: Kabir, Mohsinul, et al.
Pubblicazione: (2023)
di: Kabir, Mohsinul, et al.
Pubblicazione: (2023)
Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
DataNarrative: Automated Data-Driven Storytelling with Visualizations and Texts
di: Islam, Mohammed Saidul, et al.
Pubblicazione: (2024)
di: Islam, Mohammed Saidul, et al.
Pubblicazione: (2024)
Evolution of ReID: From Early Methods to LLM Integration
di: Bhuiyan, Amran, et al.
Pubblicazione: (2025)
di: Bhuiyan, Amran, et al.
Pubblicazione: (2025)
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2024)
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2024)
ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
di: Masry, Ahmed, et al.
Pubblicazione: (2025)
di: Masry, Ahmed, et al.
Pubblicazione: (2025)
DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
di: Kartha, Aaryaman, et al.
Pubblicazione: (2025)
di: Kartha, Aaryaman, et al.
Pubblicazione: (2025)
A Comprehensive Evaluation of Large Language Models on Benchmark Biomedical Text Processing Tasks
di: Jahan, Israt, et al.
Pubblicazione: (2023)
di: Jahan, Israt, et al.
Pubblicazione: (2023)
Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
di: Jahan, Israt, et al.
Pubblicazione: (2025)
di: Jahan, Israt, et al.
Pubblicazione: (2025)
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models
di: Alqahtani, Sawsan, et al.
Pubblicazione: (2026)
di: Alqahtani, Sawsan, et al.
Pubblicazione: (2026)
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2025)
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2025)
ChartInstruct: Instruction Tuning for Chart Comprehension and Reasoning
di: Masry, Ahmed, et al.
Pubblicazione: (2024)
di: Masry, Ahmed, et al.
Pubblicazione: (2024)
ChartGemma: Visual Instruction-tuning for Chart Reasoning in the Wild
di: Masry, Ahmed, et al.
Pubblicazione: (2024)
di: Masry, Ahmed, et al.
Pubblicazione: (2024)
XL-HeadTags: Leveraging Multimodal Retrieval Augmentation for the Multilingual Generation of News Headlines and Tags
di: Shohan, Faisal Tareque, et al.
Pubblicazione: (2024)
di: Shohan, Faisal Tareque, et al.
Pubblicazione: (2024)
Natural Language Generation for Visualizations: State of the Art, Challenges and Future Directions
di: Hoque, Enamul, et al.
Pubblicazione: (2024)
di: Hoque, Enamul, et al.
Pubblicazione: (2024)
LFOSum: Summarizing Long-form Opinions with Large Language Models
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2024)
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2024)
OpinioRAG: Towards Generating User-Centric Opinion Highlights from Large-scale Online Reviews
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2025)
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2025)
Which English Do LLMs Prefer? Triangulating Structural Bias Towards American English in Foundation Models
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2026)
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2026)
KidLM: Advancing Language Models for Children -- Early Insights and Future Directions
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2024)
di: Nayeem, Mir Tafseer, et al.
Pubblicazione: (2024)
Utilizing BERT for Information Retrieval: Survey, Applications, Resources, and Challenges
di: Wang, Jiajia, et al.
Pubblicazione: (2024)
di: Wang, Jiajia, et al.
Pubblicazione: (2024)
Open-RAG: Enhanced Retrieval-Augmented Reasoning with Open-Source Large Language Models
di: Islam, Shayekh Bin, et al.
Pubblicazione: (2024)
di: Islam, Shayekh Bin, et al.
Pubblicazione: (2024)
SurveyGen: Quality-Aware Scientific Survey Generation with Large Language Models
di: Bao, Tong, et al.
Pubblicazione: (2025)
di: Bao, Tong, et al.
Pubblicazione: (2025)
Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?
di: Fu, Xue-Yong, et al.
Pubblicazione: (2024)
di: Fu, Xue-Yong, et al.
Pubblicazione: (2024)
Position: Beyond Assistance -- Reimagining LLMs as Ethical and Adaptive Co-Creators in Mental Health Care
di: Badawi, Abeer, et al.
Pubblicazione: (2025)
di: Badawi, Abeer, et al.
Pubblicazione: (2025)
Impact of Salinisation on the Neighbour-based Spatial Diversity of Tree Species in the Sundarbans Mangrove of Bangladesh
di: Rahman, Md Mizanur
Pubblicazione: (2026)
di: Rahman, Md Mizanur
Pubblicazione: (2026)
eC-Tab2Text: Aspect-Based Text Generation from e-Commerce Product Tables
di: Guanilo, Luis Antonio Gutiérrez, et al.
Pubblicazione: (2025)
di: Guanilo, Luis Antonio Gutiérrez, et al.
Pubblicazione: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
su2-3d-lgt-metropolis: finite-volume stability and benchmark release for 3D SU(2) lattice gauge theory
di: Miraz, MD Mizanur Rahman
Pubblicazione: (2026)
di: Miraz, MD Mizanur Rahman
Pubblicazione: (2026)
Improved Detection and Diagnosis of Faults in Deep Neural Networks Using Hierarchical and Explainable Classification
di: Jahan, Sigma, et al.
Pubblicazione: (2025)
di: Jahan, Sigma, et al.
Pubblicazione: (2025)
Learning to Fast Unrank in Collaborative Filtering Recommendation
di: Zhao, Junpeng, et al.
Pubblicazione: (2025)
di: Zhao, Junpeng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Lost in Translation: Do LVLM Judges Generalize Across Languages?
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2026) -
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025) -
From Charts to Fair Narratives: Uncovering and Mitigating Geo-Economic Biases in Chart-to-Text
di: Mahbub, Ridwan, et al.
Pubblicazione: (2025) -
The Perils of Chart Deception: How Misleading Visualizations Affect Vision-Language Models
di: Mahbub, Ridwan, et al.
Pubblicazione: (2025) -
LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions
di: Rahman, Mizanur, et al.
Pubblicazione: (2025)