Automated Evaluation can Distinguish the Good and Bad AI Responses to Patient Questions about Hospitalization
Fuente:
arXiv
Saved in:
| Main Authors: | Soni, Sarvesh, Demner-Fushman, Dina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward Relieving Clinician Burden by Automatically Generating Progress Notes using Interim Hospital Data
by: Soni, Sarvesh, et al.
Published: (2024)
by: Soni, Sarvesh, et al.
Published: (2024)
A Dataset for Addressing Patient's Information Needs related to Clinical Course of Hospitalization
by: Soni, Sarvesh, et al.
Published: (2025)
by: Soni, Sarvesh, et al.
Published: (2025)
Overview of the ClinIQLink 2025 Shared Task on Medical Question-Answering
by: Colelough, Brandon, et al.
Published: (2025)
by: Colelough, Brandon, et al.
Published: (2025)
Quantifying Hallucinations in Language Language Models on Medical Textbooks
by: Colelough, Brandon C., et al.
Published: (2026)
by: Colelough, Brandon C., et al.
Published: (2026)
BioACE: An Automated Framework for Biomedical Answer and Citation Evaluations
by: Gupta, Deepak, et al.
Published: (2026)
by: Gupta, Deepak, et al.
Published: (2026)
Lessons from the TREC Plain Language Adaptation of Biomedical Abstracts (PLABA) track
by: Ondov, Brian, et al.
Published: (2025)
by: Ondov, Brian, et al.
Published: (2025)
A Dataset and Benchmark for Consumer Healthcare Question Summarization
by: Basu, Abhishek, et al.
Published: (2025)
by: Basu, Abhishek, et al.
Published: (2025)
A Dataset and Resources for Identifying Patient Health Literacy Information from Clinical Notes
by: Bittner, Madeline, et al.
Published: (2026)
by: Bittner, Madeline, et al.
Published: (2026)
Overview of TREC 2024 Medical Video Question Answering (MedVidQA) Track
by: Gupta, Deepak, et al.
Published: (2024)
by: Gupta, Deepak, et al.
Published: (2024)
Systematic Evaluation of the Quality of Synthetic Clinical Notes Rephrased by LLMs at Million-Note Scale
by: Liu, Jinghui, et al.
Published: (2026)
by: Liu, Jinghui, et al.
Published: (2026)
JEBS: A Fine-grained Biomedical Lexical Simplification Task
by: Xia, William, et al.
Published: (2025)
by: Xia, William, et al.
Published: (2025)
Reverse Probing: Supervised Token-level Uncertainty Quantification for Large Language Models in Clinical Text
by: Xiao, Bushi, et al.
Published: (2026)
by: Xiao, Bushi, et al.
Published: (2026)
The Good, The Bad, and Why: Unveiling Emotions in Generative AI
by: Li, Cheng, et al.
Published: (2023)
by: Li, Cheng, et al.
Published: (2023)
The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
by: Song, Yifan, et al.
Published: (2024)
by: Song, Yifan, et al.
Published: (2024)
Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?
by: Chen, Pinzhen, et al.
Published: (2024)
by: Chen, Pinzhen, et al.
Published: (2024)
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
by: Yi, Zihao, et al.
Published: (2025)
by: Yi, Zihao, et al.
Published: (2025)
Identifying Good and Bad Neurons for Task-Level Controllable LLMs
by: Li, Wenjie, et al.
Published: (2026)
by: Li, Wenjie, et al.
Published: (2026)
When Bad Data Leads to Good Models
by: Li, Kenneth, et al.
Published: (2025)
by: Li, Kenneth, et al.
Published: (2025)
Looks can be Deceptive: Distinguishing Repetition Disfluency from Reduplication
by: Ahmad, Arif, et al.
Published: (2024)
by: Ahmad, Arif, et al.
Published: (2024)
FADE: Why Bad Descriptions Happen to Good Features
by: Puri, Bruno, et al.
Published: (2025)
by: Puri, Bruno, et al.
Published: (2025)
The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
by: Sadallah, Abdelrahman, et al.
Published: (2025)
by: Sadallah, Abdelrahman, et al.
Published: (2025)
Can LLMs Ask Good Questions?
by: Zhang, Yueheng, et al.
Published: (2025)
by: Zhang, Yueheng, et al.
Published: (2025)
The Good, the Bad, and the Ugly: The Role of AI Quality Disclosure in Lie Detection
by: Bhattacharya, Haimanti, et al.
Published: (2024)
by: Bhattacharya, Haimanti, et al.
Published: (2024)
Agent-based Automated Claim Matching with Instruction-following LLMs
by: Pisarevskaya, Dina, et al.
Published: (2025)
by: Pisarevskaya, Dina, et al.
Published: (2025)
Small Edits, Big Consequences: Telling Good from Bad Robustness in Large Language Models
by: Ismailov, Altynbek, et al.
Published: (2025)
by: Ismailov, Altynbek, et al.
Published: (2025)
Ask Good Questions for Large Language Models
by: Wu, Qi, et al.
Published: (2025)
by: Wu, Qi, et al.
Published: (2025)
Responsible AI: The Good, The Bad, The AI
by: Jafari, Akbar Anbar, et al.
Published: (2026)
by: Jafari, Akbar Anbar, et al.
Published: (2026)
Simulation, Modelling and Classification of Wiki Contributors: Spotting The Good, The Bad, and The Ugly
by: Méndez, Silvia García, et al.
Published: (2024)
by: Méndez, Silvia García, et al.
Published: (2024)
The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG)
by: Zeng, Shenglai, et al.
Published: (2024)
by: Zeng, Shenglai, et al.
Published: (2024)
Zero-shot and Few-shot Learning with Instruction-following LLMs for Claim Matching in Automated Fact-checking
by: Pisarevskaya, Dina, et al.
Published: (2025)
by: Pisarevskaya, Dina, et al.
Published: (2025)
Rethinking KenLM: Good and Bad Model Ensembles for Efficient Text Quality Filtering in Large Web Corpora
by: Kim, Yungi, et al.
Published: (2024)
by: Kim, Yungi, et al.
Published: (2024)
Bad Actor, Good Advisor: Exploring the Role of Large Language Models in Fake News Detection
by: Hu, Beizhe, et al.
Published: (2023)
by: Hu, Beizhe, et al.
Published: (2023)
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
by: Mazumder, Aritra, et al.
Published: (2026)
by: Mazumder, Aritra, et al.
Published: (2026)
Assessing Good, Bad and Ugly Arguments Generated by ChatGPT: a New Dataset, its Methodology and Associated Tasks
by: Rocha, Victor Hugo Nascimento, et al.
Published: (2024)
by: Rocha, Victor Hugo Nascimento, et al.
Published: (2024)
MIRROR: A Novel Approach for the Automated Evaluation of Open-Ended Question Generation
by: Deroy, Aniket, et al.
Published: (2024)
by: Deroy, Aniket, et al.
Published: (2024)
How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?
by: Lo, Leo Yu-Ho, et al.
Published: (2024)
by: Lo, Leo Yu-Ho, et al.
Published: (2024)
The Future of Learning in the Age of Generative AI: Automated Question Generation and Assessment with Large Language Models
by: Maity, Subhankar, et al.
Published: (2024)
by: Maity, Subhankar, et al.
Published: (2024)
The Responsible Development of Automated Student Feedback with Generative AI
by: Lindsay, Euan D, et al.
Published: (2023)
by: Lindsay, Euan D, et al.
Published: (2023)
ASTRID -- An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems
by: Chowdhury, Mohita, et al.
Published: (2025)
by: Chowdhury, Mohita, et al.
Published: (2025)
Automated Generation of Curriculum-Aligned Multiple-Choice Questions for Malaysian Secondary Mathematics Using Generative AI
by: Wahid, Rohaizah Abdul, et al.
Published: (2025)
by: Wahid, Rohaizah Abdul, et al.
Published: (2025)
Similar Items
-
Toward Relieving Clinician Burden by Automatically Generating Progress Notes using Interim Hospital Data
by: Soni, Sarvesh, et al.
Published: (2024) -
A Dataset for Addressing Patient's Information Needs related to Clinical Course of Hospitalization
by: Soni, Sarvesh, et al.
Published: (2025) -
Overview of the ClinIQLink 2025 Shared Task on Medical Question-Answering
by: Colelough, Brandon, et al.
Published: (2025) -
Quantifying Hallucinations in Language Language Models on Medical Textbooks
by: Colelough, Brandon C., et al.
Published: (2026) -
BioACE: An Automated Framework for Biomedical Answer and Citation Evaluations
by: Gupta, Deepak, et al.
Published: (2026)