Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions
Fuente:
arXiv
Saved in:
| Main Authors: | Raman, Naveen, Cortes-Gomez, Santiago, Rubio, Mateo Dulce, Fang, Fei, Wilder, Bryan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Global Rewards in Restless Multi-Armed Bandits
by: Raman, Naveen, et al.
Published: (2024)
by: Raman, Naveen, et al.
Published: (2024)
Auditing Fairness by Betting
by: Chugg, Ben, et al.
Published: (2023)
by: Chugg, Ben, et al.
Published: (2023)
Contextual Budget Bandit for Food Rescue Volunteer Engagement
by: Tang, Ariana, et al.
Published: (2025)
by: Tang, Ariana, et al.
Published: (2025)
Data-driven Design of Randomized Control Trials with Guaranteed Treatment Effects
by: Cortes-Gomez, Santiago, et al.
Published: (2024)
by: Cortes-Gomez, Santiago, et al.
Published: (2024)
The Limits of AI-Driven Allocation: Optimal Screening under Aleatoric Uncertainty
by: Cortes-Gomez, Santiago, et al.
Published: (2026)
by: Cortes-Gomez, Santiago, et al.
Published: (2026)
Selecting Decision-Relevant Concepts in Reinforcement Learning
by: Raman, Naveen, et al.
Published: (2026)
by: Raman, Naveen, et al.
Published: (2026)
Transformers in Healthcare: A Survey
by: Nerella, Subhash, et al.
Published: (2023)
by: Nerella, Subhash, et al.
Published: (2023)
Predicting Healthcare Provider Engagement in SMS Campaigns
by: Qureshi, Daanish Aleem, et al.
Published: (2025)
by: Qureshi, Daanish Aleem, et al.
Published: (2025)
An Epistemic and Aleatoric Decomposition of Arbitrariness to Constrain the Set of Good Models
by: Khan, Falaah Arif, et al.
Published: (2023)
by: Khan, Falaah Arif, et al.
Published: (2023)
Fair Machine Learning in Healthcare: A Review
by: Feng, Qizhang, et al.
Published: (2022)
by: Feng, Qizhang, et al.
Published: (2022)
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
by: Potham, Ram
Published: (2025)
by: Potham, Ram
Published: (2025)
Generative Artificial Intelligence in Healthcare: Ethical Considerations and Assessment Checklist
by: Ning, Yilin, et al.
Published: (2023)
by: Ning, Yilin, et al.
Published: (2023)
United States Road Accident Prediction using Random Forest Predictor
by: Yamarthi, Dominic Parosh, et al.
Published: (2025)
by: Yamarthi, Dominic Parosh, et al.
Published: (2025)
RescueLens: LLM-Powered Triage and Action on Volunteer Feedback for Food Rescue
by: Raman, Naveen, et al.
Published: (2025)
by: Raman, Naveen, et al.
Published: (2025)
Not All Options Are Created Equal: Textual Option Weighting for Token-Efficient LLM-Based Knowledge Tracing
by: Kim, JongWoo, et al.
Published: (2024)
by: Kim, JongWoo, et al.
Published: (2024)
Fairness-Aware Interpretable Modeling (FAIM) for Trustworthy Machine Learning in Healthcare
by: Liu, Mingxuan, et al.
Published: (2024)
by: Liu, Mingxuan, et al.
Published: (2024)
Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare
by: Tariq, Amara, et al.
Published: (2025)
by: Tariq, Amara, et al.
Published: (2025)
Implementation of Big Data Analytics for Diabetes Management: Needs Assessment in the Rwanda Healthcare System
by: Majyambere, Silas, et al.
Published: (2026)
by: Majyambere, Silas, et al.
Published: (2026)
Integrating Social Determinants of Health into Knowledge Graphs: Evaluating Prediction Bias and Fairness in Healthcare
by: Shang, Tianqi, et al.
Published: (2024)
by: Shang, Tianqi, et al.
Published: (2024)
FairHealth: An Open-Source Python Library for Trustworthy Healthcare AI in Low-Resource Settings
by: Yesmin, Farjana
Published: (2026)
by: Yesmin, Farjana
Published: (2026)
Do Concept Bottleneck Models Respect Localities?
by: Raman, Naveen, et al.
Published: (2024)
by: Raman, Naveen, et al.
Published: (2024)
Fostering the Ecosystem of AI for Social Impact Requires Expanding and Strengthening Evaluation Standards
by: Wilder, Bryan, et al.
Published: (2025)
by: Wilder, Bryan, et al.
Published: (2025)
Deprecating Benchmarks: Criteria and Framework
by: Joaquin, Ayrton San, et al.
Published: (2025)
by: Joaquin, Ayrton San, et al.
Published: (2025)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
by: Guldimann, Philipp, et al.
Published: (2024)
by: Guldimann, Philipp, et al.
Published: (2024)
Recommender Systems for Good (RS4Good): Survey of Use Cases and a Call to Action for Research that Matters
by: Jannach, Dietmar, et al.
Published: (2024)
by: Jannach, Dietmar, et al.
Published: (2024)
Moral Alignment for LLM Agents
by: Tennant, Elizaveta, et al.
Published: (2024)
by: Tennant, Elizaveta, et al.
Published: (2024)
The Odyssey of the Fittest: Can Agents Survive and Still Be Good?
by: Waldner, Dylan, et al.
Published: (2025)
by: Waldner, Dylan, et al.
Published: (2025)
LLM Safety Alignment is Divergence Estimation in Disguise
by: Haldar, Rajdeep, et al.
Published: (2025)
by: Haldar, Rajdeep, et al.
Published: (2025)
Selecting the Right LLM for eGov Explanations
by: Limonad, Lior, et al.
Published: (2025)
by: Limonad, Lior, et al.
Published: (2025)
Mapping the Media Landscape: Predicting Factual Reporting and Political Bias Through Web Interactions
by: Sánchez-Cortés, Dairazalia, et al.
Published: (2024)
by: Sánchez-Cortés, Dairazalia, et al.
Published: (2024)
FFB: A Fair Fairness Benchmark for In-Processing Group Fairness Methods
by: Han, Xiaotian, et al.
Published: (2023)
by: Han, Xiaotian, et al.
Published: (2023)
DM-Bench: Benchmarking LLMs for Personalized Decision Making in Diabetes Management
by: Cardei, Maria Ana, et al.
Published: (2025)
by: Cardei, Maria Ana, et al.
Published: (2025)
TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting Methods
by: Qiu, Xiangfei, et al.
Published: (2024)
by: Qiu, Xiangfei, et al.
Published: (2024)
TokenPowerBench: Benchmarking the Power Consumption of LLM Inference
by: Niu, Chenxu, et al.
Published: (2025)
by: Niu, Chenxu, et al.
Published: (2025)
Scaling Legal AI: Benchmarking Mamba and Transformers for Statutory Classification and Case Law Retrieval
by: Maurya, Anuraj
Published: (2025)
by: Maurya, Anuraj
Published: (2025)
From Protoscience to Epistemic Monoculture: How Benchmarking Set the Stage for the Deep Learning Revolution
by: Koch, Bernard J., et al.
Published: (2024)
by: Koch, Bernard J., et al.
Published: (2024)
NaiAD: Initiate Data-Driven Research for LLM Advertising
by: Zhang, Yihang, et al.
Published: (2026)
by: Zhang, Yihang, et al.
Published: (2026)
When the Domain Expert Has No Time and the LLM Developer Has No Clinical Expertise: Real-World Lessons from LLM Co-Design in a Safety-Net Hospital
by: Kothari, Avni, et al.
Published: (2025)
by: Kothari, Avni, et al.
Published: (2025)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture
by: Ackerman, Gary, et al.
Published: (2025)
by: Ackerman, Gary, et al.
Published: (2025)
A Comparative Benchmark of Federated Learning Strategies for Mortality Prediction on Heterogeneous and Imbalanced Clinical Data
by: Tertulino, Rodrigo
Published: (2025)
by: Tertulino, Rodrigo
Published: (2025)
Similar Items
-
Global Rewards in Restless Multi-Armed Bandits
by: Raman, Naveen, et al.
Published: (2024) -
Auditing Fairness by Betting
by: Chugg, Ben, et al.
Published: (2023) -
Contextual Budget Bandit for Food Rescue Volunteer Engagement
by: Tang, Ariana, et al.
Published: (2025) -
Data-driven Design of Randomized Control Trials with Guaranteed Treatment Effects
by: Cortes-Gomez, Santiago, et al.
Published: (2024) -
The Limits of AI-Driven Allocation: Optimal Screening under Aleatoric Uncertainty
by: Cortes-Gomez, Santiago, et al.
Published: (2026)