Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge
Fuente:
arXiv
Saved in:
| Main Authors: | Laskar, Md Tahmid Rahman, Jahan, Israt, Dolatabadi, Elham, Peng, Chun, Hoque, Enamul, Huang, Jimmy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
by: Jahan, Israt, et al.
Published: (2025)
by: Jahan, Israt, et al.
Published: (2025)
A Comprehensive Evaluation of Large Language Models on Benchmark Biomedical Text Processing Tasks
by: Jahan, Israt, et al.
Published: (2023)
by: Jahan, Israt, et al.
Published: (2023)
Position: Beyond Assistance -- Reimagining LLMs as Ethical and Adaptive Co-Creators in Mental Health Care
by: Badawi, Abeer, et al.
Published: (2025)
by: Badawi, Abeer, et al.
Published: (2025)
Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
by: Badawi, Abeer, et al.
Published: (2025)
by: Badawi, Abeer, et al.
Published: (2025)
Lost in Translation: Do LVLM Judges Generalize Across Languages?
by: Laskar, Md Tahmid Rahman, et al.
Published: (2026)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2026)
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
by: Rahman, Mizanur, et al.
Published: (2025)
by: Rahman, Mizanur, et al.
Published: (2025)
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations
by: Laskar, Md Tahmid Rahman, et al.
Published: (2024)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2024)
Are Large Vision Language Models up to the Challenge of Chart Comprehension and Reasoning? An Extensive Investigation into the Capabilities and Limitations of LVLMs
by: Islam, Mohammed Saidul, et al.
Published: (2024)
by: Islam, Mohammed Saidul, et al.
Published: (2024)
BenLLMEval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP
by: Kabir, Mohsinul, et al.
Published: (2023)
by: Kabir, Mohsinul, et al.
Published: (2023)
Aligning Text, Code, and Vision: A Multi-Objective Reinforcement Learning Framework for Text-to-Visualization
by: Rahman, Mizanur, et al.
Published: (2026)
by: Rahman, Mizanur, et al.
Published: (2026)
DataNarrative: Automated Data-Driven Storytelling with Visualizations and Texts
by: Islam, Mohammed Saidul, et al.
Published: (2024)
by: Islam, Mohammed Saidul, et al.
Published: (2024)
The Perils of Chart Deception: How Misleading Visualizations Affect Vision-Language Models
by: Mahbub, Ridwan, et al.
Published: (2025)
by: Mahbub, Ridwan, et al.
Published: (2025)
Assessing the Quality of Mental Health Support in LLM Responses through Multi-Attribute Human Evaluation
by: Badawi, Abeer, et al.
Published: (2026)
by: Badawi, Abeer, et al.
Published: (2026)
From Charts to Fair Narratives: Uncovering and Mitigating Geo-Economic Biases in Chart-to-Text
by: Mahbub, Ridwan, et al.
Published: (2025)
by: Mahbub, Ridwan, et al.
Published: (2025)
Evolution of ReID: From Early Methods to LLM Integration
by: Bhuiyan, Amran, et al.
Published: (2025)
by: Bhuiyan, Amran, et al.
Published: (2025)
LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions
by: Rahman, Mizanur, et al.
Published: (2025)
by: Rahman, Mizanur, et al.
Published: (2025)
Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages
by: Alam, Firoj, et al.
Published: (2026)
by: Alam, Firoj, et al.
Published: (2026)
Query-OPT: Optimizing Inference of Large Language Models via Multi-Query Instructions in Meeting Summarization
by: Laskar, Md Tahmid Rahman, et al.
Published: (2024)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2024)
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models
by: Alqahtani, Sawsan, et al.
Published: (2026)
by: Alqahtani, Sawsan, et al.
Published: (2026)
DACP: Domain-Adaptive Continual Pre-Training of Large Language Models for Phone Conversation Summarization
by: Fu, Xue-Yong, et al.
Published: (2025)
by: Fu, Xue-Yong, et al.
Published: (2025)
Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?
by: Fu, Xue-Yong, et al.
Published: (2024)
by: Fu, Xue-Yong, et al.
Published: (2024)
Natural Language Generation for Visualizations: State of the Art, Challenges and Future Directions
by: Hoque, Enamul, et al.
Published: (2024)
by: Hoque, Enamul, et al.
Published: (2024)
Open-RAG: Enhanced Retrieval-Augmented Reasoning with Open-Source Large Language Models
by: Islam, Shayekh Bin, et al.
Published: (2024)
by: Islam, Shayekh Bin, et al.
Published: (2024)
RELATE: Relation Extraction in Biomedical Abstracts with LLMs and Ontology Constraints
by: Olasunkanmi, Olawumi, et al.
Published: (2025)
by: Olasunkanmi, Olawumi, et al.
Published: (2025)
Utilizing BERT for Information Retrieval: Survey, Applications, Resources, and Challenges
by: Wang, Jiajia, et al.
Published: (2024)
by: Wang, Jiajia, et al.
Published: (2024)
ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
User Profile with Large Language Models: Construction, Updating, and Benchmarking
by: Prottasha, Nusrat Jahan, et al.
Published: (2025)
by: Prottasha, Nusrat Jahan, et al.
Published: (2025)
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
by: Nayeem, Mir Tafseer, et al.
Published: (2025)
by: Nayeem, Mir Tafseer, et al.
Published: (2025)
Decision-Aware Trust Signal Alignment for SOC Alert Triage
by: Chowdhury, Israt Jahan, et al.
Published: (2026)
by: Chowdhury, Israt Jahan, et al.
Published: (2026)
IoTWarden: A Deep Reinforcement Learning Based Real-time Defense System to Mitigate Trigger-action IoT Attacks
by: Alam, Md Morshed, et al.
Published: (2024)
by: Alam, Md Morshed, et al.
Published: (2024)
A Flexible Fairness Framework with Surrogate Loss Reweighting for Addressing Sociodemographic Disparities
by: Xu, Wen, et al.
Published: (2025)
by: Xu, Wen, et al.
Published: (2025)
Tangail Saree as Geographical Indication of Bangladesh
by: Efat, Shanjida Israt Jahan
Published: (2025)
by: Efat, Shanjida Israt Jahan
Published: (2025)
Fine-tuned Large Language Models (LLMs): Improved Prompt Injection Attacks Detection
by: Rahman, Md Abdur, et al.
Published: (2024)
by: Rahman, Md Abdur, et al.
Published: (2024)
A Heterogeneous Two-Stream Framework for Video Action Recognition with Comparative Fusion Analysis
by: Rahaman, Md. Afzalur, et al.
Published: (2026)
by: Rahaman, Md. Afzalur, et al.
Published: (2026)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
by: Thakur, Aman Singh, et al.
Published: (2024)
by: Thakur, Aman Singh, et al.
Published: (2024)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
by: Badshah, Sher, et al.
Published: (2024)
by: Badshah, Sher, et al.
Published: (2024)
Enhancing Salt Tolerance in Foxtail Millet Through Organic Inputs: A Soil–Plant–Microbe Interaction Perspective
by: Israt Jahan Irin, et al.
Published: (2026)
by: Israt Jahan Irin, et al.
Published: (2026)
Enhancing Financial Knowledge Through High School Education: The Effect of Mandated Economics and Personal Finance Courses
by: Taufiq Hasan Quadria, et al.
Published: (2025)
by: Taufiq Hasan Quadria, et al.
Published: (2025)
Similar Items
-
Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
by: Jahan, Israt, et al.
Published: (2025) -
A Comprehensive Evaluation of Large Language Models on Benchmark Biomedical Text Processing Tasks
by: Jahan, Israt, et al.
Published: (2023) -
Position: Beyond Assistance -- Reimagining LLMs as Ethical and Adaptive Co-Creators in Mental Health Care
by: Badawi, Abeer, et al.
Published: (2025) -
Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025) -
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)