Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jahan, Israt, Laskar, Md Tahmid Rahman, Peng, Chun, Huang, Jimmy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Comprehensive Evaluation of Large Language Models on Benchmark Biomedical Text Processing Tasks
von: Jahan, Israt, et al.
Veröffentlicht: (2023)
von: Jahan, Israt, et al.
Veröffentlicht: (2023)
Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2024)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2024)
User Profile with Large Language Models: Construction, Updating, and Benchmarking
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2025)
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2025)
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
von: Badawi, Abeer, et al.
Veröffentlicht: (2025)
von: Badawi, Abeer, et al.
Veröffentlicht: (2025)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
BenLLMEval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2023)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2023)
Query-OPT: Optimizing Inference of Large Language Models via Multi-Query Instructions in Meeting Summarization
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2024)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2024)
Are Large Vision Language Models up to the Challenge of Chart Comprehension and Reasoning? An Extensive Investigation into the Capabilities and Limitations of LVLMs
von: Islam, Mohammed Saidul, et al.
Veröffentlicht: (2024)
von: Islam, Mohammed Saidul, et al.
Veröffentlicht: (2024)
Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?
von: Fu, Xue-Yong, et al.
Veröffentlicht: (2024)
von: Fu, Xue-Yong, et al.
Veröffentlicht: (2024)
Lost in Translation: Do LVLM Judges Generalize Across Languages?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2026)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2026)
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models
von: Alqahtani, Sawsan, et al.
Veröffentlicht: (2026)
von: Alqahtani, Sawsan, et al.
Veröffentlicht: (2026)
The Perils of Chart Deception: How Misleading Visualizations Affect Vision-Language Models
von: Mahbub, Ridwan, et al.
Veröffentlicht: (2025)
von: Mahbub, Ridwan, et al.
Veröffentlicht: (2025)
DACP: Domain-Adaptive Continual Pre-Training of Large Language Models for Phone Conversation Summarization
von: Fu, Xue-Yong, et al.
Veröffentlicht: (2025)
von: Fu, Xue-Yong, et al.
Veröffentlicht: (2025)
Aligning Text, Code, and Vision: A Multi-Objective Reinforcement Learning Framework for Text-to-Visualization
von: Rahman, Mizanur, et al.
Veröffentlicht: (2026)
von: Rahman, Mizanur, et al.
Veröffentlicht: (2026)
DataNarrative: Automated Data-Driven Storytelling with Visualizations and Texts
von: Islam, Mohammed Saidul, et al.
Veröffentlicht: (2024)
von: Islam, Mohammed Saidul, et al.
Veröffentlicht: (2024)
Utilizing BERT for Information Retrieval: Survey, Applications, Resources, and Challenges
von: Wang, Jiajia, et al.
Veröffentlicht: (2024)
von: Wang, Jiajia, et al.
Veröffentlicht: (2024)
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models
von: Jiang, Junfeng, et al.
Veröffentlicht: (2024)
von: Jiang, Junfeng, et al.
Veröffentlicht: (2024)
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
von: Nayeem, Mir Tafseer, et al.
Veröffentlicht: (2025)
von: Nayeem, Mir Tafseer, et al.
Veröffentlicht: (2025)
Evolution of ReID: From Early Methods to LLM Integration
von: Bhuiyan, Amran, et al.
Veröffentlicht: (2025)
von: Bhuiyan, Amran, et al.
Veröffentlicht: (2025)
From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2026)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2026)
From Charts to Fair Narratives: Uncovering and Mitigating Geo-Economic Biases in Chart-to-Text
von: Mahbub, Ridwan, et al.
Veröffentlicht: (2025)
von: Mahbub, Ridwan, et al.
Veröffentlicht: (2025)
ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)
Large Language Models for IT Automation Tasks: Are We There Yet?
von: Hassan, Md Mahadi, et al.
Veröffentlicht: (2025)
von: Hassan, Md Mahadi, et al.
Veröffentlicht: (2025)
DACIP-RC: Domain Adaptive Continual Instruction Pre-Training via Reading Comprehension on Business Conversations
von: Khasanova, Elena, et al.
Veröffentlicht: (2025)
von: Khasanova, Elena, et al.
Veröffentlicht: (2025)
Parameter-Efficient Fine-Tuning of Large Language Models using Semantic Knowledge Tuning
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2024)
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2024)
AI Knowledge Assist: An Automated Approach for the Creation of Knowledge Bases for Conversational AI Agents
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models on Spatial Tasks: A Multi-Task Benchmarking Study
von: Xu, Liuchang, et al.
Veröffentlicht: (2024)
von: Xu, Liuchang, et al.
Veröffentlicht: (2024)
Monkey Jump : MoE-Style PEFT for Efficient Multi-Task Learning
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2026)
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2026)
Benchmarking Retrieval-Augmented Large Language Models in Biomedical NLP: Application, Robustness, and Self-Awareness
von: Li, Mingchen, et al.
Veröffentlicht: (2024)
von: Li, Mingchen, et al.
Veröffentlicht: (2024)
Position: Beyond Assistance -- Reimagining LLMs as Ethical and Adaptive Co-Creators in Mental Health Care
von: Badawi, Abeer, et al.
Veröffentlicht: (2025)
von: Badawi, Abeer, et al.
Veröffentlicht: (2025)
Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation
von: Zhu, Qin, et al.
Veröffentlicht: (2024)
von: Zhu, Qin, et al.
Veröffentlicht: (2024)
Large Language Model Benchmarks in Medical Tasks
von: Yan, Lawrence K. Q., et al.
Veröffentlicht: (2024)
von: Yan, Lawrence K. Q., et al.
Veröffentlicht: (2024)
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
Large Language Models as Biomedical Hypothesis Generators: A Comprehensive Evaluation
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
PM-LLM-Benchmark: Evaluating Large Language Models on Process Mining Tasks
von: Berti, Alessandro, et al.
Veröffentlicht: (2024)
von: Berti, Alessandro, et al.
Veröffentlicht: (2024)
Pitfalls of Evaluating Language Models with Open Benchmarks
von: Hasan, Md. Najib, et al.
Veröffentlicht: (2025)
von: Hasan, Md. Najib, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Comprehensive Evaluation of Large Language Models on Benchmark Biomedical Text Processing Tasks
von: Jahan, Israt, et al.
Veröffentlicht: (2023) -
Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025) -
Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025) -
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2024) -
User Profile with Large Language Models: Construction, Updating, and Benchmarking
von: Prottasha, Nusrat Jahan, et al.
Veröffentlicht: (2025)