Evaluating LLMs in Medicine: A Call for Rigor, Transparency
Fuente:
arXiv
Saved in:
| Main Authors: | Alwakeel, Mahmoud, Nagori, Aditya, Krishnamoorthy, Vijay, Kamaleswaran, Rishikesan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Performance of Large Language Models in Answering Critical Care Medicine Questions
by: Alwakeel, Mahmoud, et al.
Published: (2025)
by: Alwakeel, Mahmoud, et al.
Published: (2025)
Planner-Auditor Twin: Agentic Discharge Planning with FHIR-Based LLM Planning, Guideline Recall, Optional Caching and Self-Improvement
by: Wu, Kaiyuan, et al.
Published: (2026)
by: Wu, Kaiyuan, et al.
Published: (2026)
Contextual Phenotyping of Pediatric Sepsis Cohort Using Large Language Models
by: Nagori, Aditya, et al.
Published: (2025)
by: Nagori, Aditya, et al.
Published: (2025)
Open-Source Agentic Hybrid RAG Framework for Scientific Literature Review
by: Nagori, Aditya, et al.
Published: (2025)
by: Nagori, Aditya, et al.
Published: (2025)
Social Media as a Sensor: Analyzing Twitter Data for Breast Cancer Medication Effects Using Natural Language Processing
by: Kobara, Seibi, et al.
Published: (2024)
by: Kobara, Seibi, et al.
Published: (2024)
SpikGPT: A High-Accuracy and Interpretable Spiking Attention Framework for Single-Cell Annotation
by: Huang, Min, et al.
Published: (2025)
by: Huang, Min, et al.
Published: (2025)
CXR-TFT: Multi-Modal Temporal Fusion Transformer for Predicting Chest X-ray Trajectories
by: Arora, Mehak, et al.
Published: (2025)
by: Arora, Mehak, et al.
Published: (2025)
Call for Rigor in Reporting Quality of Instruction Tuning Data
by: Moon, Hyeonseok, et al.
Published: (2025)
by: Moon, Hyeonseok, et al.
Published: (2025)
Detecting algorithmic bias in medical-AI models using trees
by: Smith, Jeffrey, et al.
Published: (2023)
by: Smith, Jeffrey, et al.
Published: (2023)
Predicting Antibiotic Resistance Patterns Using Sentence-BERT: A Machine Learning Approach
by: Alwakeel, Mahmoud, et al.
Published: (2025)
by: Alwakeel, Mahmoud, et al.
Published: (2025)
NeuroSep-CP-LCB: A Deep Learning-based Contextual Multi-armed Bandit Algorithm with Uncertainty Quantification for Early Sepsis Prediction
by: Zhou, Anni, et al.
Published: (2025)
by: Zhou, Anni, et al.
Published: (2025)
Evaluating Large Language Models for Automated Clinical Abstraction in Pulmonary Embolism Registries: Performance Across Model Sizes, Versions, and Parameters
by: Alwakeel, Mahmoud, et al.
Published: (2025)
by: Alwakeel, Mahmoud, et al.
Published: (2025)
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
by: Basu, Kinjal, et al.
Published: (2024)
by: Basu, Kinjal, et al.
Published: (2024)
IslamicMMLU: A Benchmark for Evaluating LLMs on Islamic Knowledge
by: Abdelaal, Ali, et al.
Published: (2026)
by: Abdelaal, Ali, et al.
Published: (2026)
Beyond Memorization: A Rigorous Evaluation Framework for Medical Knowledge Editing
by: Chen, Shigeng, et al.
Published: (2025)
by: Chen, Shigeng, et al.
Published: (2025)
A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
by: Srewa, Mahmoud, et al.
Published: (2025)
by: Srewa, Mahmoud, et al.
Published: (2025)
A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages
by: Anikina, Tatiana, et al.
Published: (2025)
by: Anikina, Tatiana, et al.
Published: (2025)
Evaluating LLMs on Sequential API Call Through Automated Test Generation
by: Huang, Yuheng, et al.
Published: (2025)
by: Huang, Yuheng, et al.
Published: (2025)
FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
by: Zhu, Fengbin, et al.
Published: (2025)
by: Zhu, Fengbin, et al.
Published: (2025)
Culture is Everywhere: A Call for Intentionally Cultural Evaluation
by: Oh, Juhyun, et al.
Published: (2025)
by: Oh, Juhyun, et al.
Published: (2025)
Tool Calling for Arabic LLMs: Data Strategies and Instruction Tuning
by: Ersoy, Asim, et al.
Published: (2025)
by: Ersoy, Asim, et al.
Published: (2025)
Are LLMs Rigorous Logical Reasoners? Empowering Natural Language Proof Generation by Stepwise Decoding with Contrastive Learning
by: Su, Ying, et al.
Published: (2023)
by: Su, Ying, et al.
Published: (2023)
AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
by: Shu, Yiheng, et al.
Published: (2026)
by: Shu, Yiheng, et al.
Published: (2026)
Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency
by: Maslenkova, Svetlana, et al.
Published: (2025)
by: Maslenkova, Svetlana, et al.
Published: (2025)
Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions
by: Hong, Pengfei, et al.
Published: (2024)
by: Hong, Pengfei, et al.
Published: (2024)
Better Call Claude: Can LLMs Detect Changes of Writing Style?
by: Römisch, Johannes, et al.
Published: (2025)
by: Römisch, Johannes, et al.
Published: (2025)
State of What Art? A Call for Multi-Prompt LLM Evaluation
by: Mizrahi, Moran, et al.
Published: (2023)
by: Mizrahi, Moran, et al.
Published: (2023)
Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning
by: Chen, Mingyang, et al.
Published: (2024)
by: Chen, Mingyang, et al.
Published: (2024)
Alopex: A Computational Framework for Enabling On-Device Function Calls with LLMs
by: Ran, Yide, et al.
Published: (2024)
by: Ran, Yide, et al.
Published: (2024)
Power-of-Two Quantization-Aware-Training (PoT-QAT) in Large Language Models (LLMs)
by: Elgenedy, Mahmoud
Published: (2026)
by: Elgenedy, Mahmoud
Published: (2026)
Sepsyn-OLCP: An Online Learning-based Framework for Early Sepsis Prediction with Uncertainty Quantification using Conformal Prediction
by: Zhou, Anni, et al.
Published: (2025)
by: Zhou, Anni, et al.
Published: (2025)
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
by: Myrzakhan, Aidar, et al.
Published: (2024)
by: Myrzakhan, Aidar, et al.
Published: (2024)
Evaluating ChatGPT as a Recommender System: A Rigorous Approach
by: Di Palma, Dario, et al.
Published: (2023)
by: Di Palma, Dario, et al.
Published: (2023)
What Makes Math Word Problems Challenging for LLMs?
by: Srivatsa, KV Aditya, et al.
Published: (2024)
by: Srivatsa, KV Aditya, et al.
Published: (2024)
MedArena: Comparing LLMs for Medicine-in-the-Wild Clinician Preferences
by: Wu, Eric, et al.
Published: (2026)
by: Wu, Eric, et al.
Published: (2026)
[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs
by: Rao, Abhinav, et al.
Published: (2024)
by: Rao, Abhinav, et al.
Published: (2024)
K-SENSE: A Knowledge-Guided Self-Augmented Encoder for Neuro-Semantic Evaluation of Mental Health Conditions on Social Media
by: Yadav, Vijay
Published: (2026)
by: Yadav, Vijay
Published: (2026)
Con-ReCall: Detecting Pre-training Data in LLMs via Contrastive Decoding
by: Wang, Cheng, et al.
Published: (2024)
by: Wang, Cheng, et al.
Published: (2024)
MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine
by: Kong, Shufeng, et al.
Published: (2025)
by: Kong, Shufeng, et al.
Published: (2025)
When2Call: When (not) to Call Tools
by: Ross, Hayley, et al.
Published: (2025)
by: Ross, Hayley, et al.
Published: (2025)
Similar Items
-
Performance of Large Language Models in Answering Critical Care Medicine Questions
by: Alwakeel, Mahmoud, et al.
Published: (2025) -
Planner-Auditor Twin: Agentic Discharge Planning with FHIR-Based LLM Planning, Guideline Recall, Optional Caching and Self-Improvement
by: Wu, Kaiyuan, et al.
Published: (2026) -
Contextual Phenotyping of Pediatric Sepsis Cohort Using Large Language Models
by: Nagori, Aditya, et al.
Published: (2025) -
Open-Source Agentic Hybrid RAG Framework for Scientific Literature Review
by: Nagori, Aditya, et al.
Published: (2025) -
Social Media as a Sensor: Analyzing Twitter Data for Breast Cancer Medication Effects Using Natural Language Processing
by: Kobara, Seibi, et al.
Published: (2024)