Beyond Accuracy: Risk-Sensitive Evaluation of Hallucinated Medical Advice
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Doshi, Savan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hallucinations and Truth: A Comprehensive Accuracy Evaluation of RAG, LoRA and DoRA
von: Baqar, Mohammad, et al.
Veröffentlicht: (2025)
von: Baqar, Mohammad, et al.
Veröffentlicht: (2025)
Beyond Facts: Evaluating Intent Hallucination in Large Language Models
von: Hao, Yijie, et al.
Veröffentlicht: (2025)
von: Hao, Yijie, et al.
Veröffentlicht: (2025)
Luna: An Evaluation Foundation Model to Catch Language Model Hallucinations with High Accuracy and Low Cost
von: Belyi, Masha, et al.
Veröffentlicht: (2024)
von: Belyi, Masha, et al.
Veröffentlicht: (2024)
Beyond Sentiment: A Multi-Agent Pipeline for Actionable Business Advice from Reviews
von: Bhandari, Kartikey Singh, et al.
Veröffentlicht: (2026)
von: Bhandari, Kartikey Singh, et al.
Veröffentlicht: (2026)
Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI
von: Li, Zihao, et al.
Veröffentlicht: (2025)
von: Li, Zihao, et al.
Veröffentlicht: (2025)
Customized Information and Domain-centric Knowledge Graph Construction with Large Language Models
von: Wawrzik, Frank, et al.
Veröffentlicht: (2024)
von: Wawrzik, Frank, et al.
Veröffentlicht: (2024)
Measuring Stability Beyond Accuracy in Small Open-Source Medical Large Language Models for Pediatric Endocrinology
von: D'Amario, Vanessa, et al.
Veröffentlicht: (2025)
von: D'Amario, Vanessa, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Evaluating the Feasibility and Accuracy of Large Language Models for Medical History-Taking in Obstetrics and Gynecology
von: Liu, Dou, et al.
Veröffentlicht: (2025)
von: Liu, Dou, et al.
Veröffentlicht: (2025)
MedHal: An Evaluation Dataset for Medical Hallucination Detection
von: Mehenni, Gaya, et al.
Veröffentlicht: (2025)
von: Mehenni, Gaya, et al.
Veröffentlicht: (2025)
Quantifying Hallucinations in Language Language Models on Medical Textbooks
von: Colelough, Brandon C., et al.
Veröffentlicht: (2026)
von: Colelough, Brandon C., et al.
Veröffentlicht: (2026)
When Chain-of-Thought Backfires: Evaluating Prompt Sensitivity in Medical Language Models
von: Sadanandan, Binesh, et al.
Veröffentlicht: (2026)
von: Sadanandan, Binesh, et al.
Veröffentlicht: (2026)
Medical Hallucinations in Foundation Models and Their Impact on Healthcare
von: Kim, Yubin, et al.
Veröffentlicht: (2025)
von: Kim, Yubin, et al.
Veröffentlicht: (2025)
This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
von: Yun, Hye Sun, et al.
Veröffentlicht: (2026)
von: Yun, Hye Sun, et al.
Veröffentlicht: (2026)
Beyond Length: Context-Aware Expansion and Independence as Developmentally Sensitive Evaluation in Child Utterances
von: Chun, Jiyun, et al.
Veröffentlicht: (2026)
von: Chun, Jiyun, et al.
Veröffentlicht: (2026)
Lynx: An Open Source Hallucination Evaluation Model
von: Ravi, Selvan Sunitha, et al.
Veröffentlicht: (2024)
von: Ravi, Selvan Sunitha, et al.
Veröffentlicht: (2024)
Beyond Accuracy: An Explainability-Driven Analysis of Harmful Content Detection
von: Dhara, Trishita, et al.
Veröffentlicht: (2026)
von: Dhara, Trishita, et al.
Veröffentlicht: (2026)
Writing in Symbiosis: Mapping Human Creative Agency in the AI Era
von: Doshi, Vivan, et al.
Veröffentlicht: (2025)
von: Doshi, Vivan, et al.
Veröffentlicht: (2025)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
von: Naderi, Nariman, et al.
Veröffentlicht: (2025)
von: Naderi, Nariman, et al.
Veröffentlicht: (2025)
Backprompting: Leveraging Synthetic Production Data for Health Advice Guardrails
von: Cheng, Kellen Tan, et al.
Veröffentlicht: (2025)
von: Cheng, Kellen Tan, et al.
Veröffentlicht: (2025)
Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Introducing a Symbolic-Mechanistic Approach to Interpretable Evaluation
von: Habibi, Reza, et al.
Veröffentlicht: (2026)
von: Habibi, Reza, et al.
Veröffentlicht: (2026)
On the Relation between Sensitivity and Accuracy in In-context Learning
von: Chen, Yanda, et al.
Veröffentlicht: (2022)
von: Chen, Yanda, et al.
Veröffentlicht: (2022)
Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
von: Huang, Liangjie, et al.
Veröffentlicht: (2025)
von: Huang, Liangjie, et al.
Veröffentlicht: (2025)
Reducing Hallucinations of Medical Multimodal Large Language Models with Visual Retrieval-Augmented Generation
von: Chu, Yun-Wei, et al.
Veröffentlicht: (2025)
von: Chu, Yun-Wei, et al.
Veröffentlicht: (2025)
Hallucination Detection and Hallucination Mitigation: An Investigation
von: Luo, Junliang, et al.
Veröffentlicht: (2024)
von: Luo, Junliang, et al.
Veröffentlicht: (2024)
A Survey of Automatic Hallucination Evaluation on Natural Language Generation
von: Qi, Siya, et al.
Veröffentlicht: (2024)
von: Qi, Siya, et al.
Veröffentlicht: (2024)
MedHallBench: A New Benchmark for Assessing Hallucination in Medical Large Language Models
von: Zuo, Kaiwen, et al.
Veröffentlicht: (2024)
von: Zuo, Kaiwen, et al.
Veröffentlicht: (2024)
Hallucination Benchmark in Medical Visual Question Answering
von: Wu, Jinge, et al.
Veröffentlicht: (2024)
von: Wu, Jinge, et al.
Veröffentlicht: (2024)
Beyond Fine-Tuning: Effective Strategies for Mitigating Hallucinations in Large Language Models for Data Analytics
von: Rumiantsau, Mikhail, et al.
Veröffentlicht: (2024)
von: Rumiantsau, Mikhail, et al.
Veröffentlicht: (2024)
Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
Removal of Hallucination on Hallucination: Debate-Augmented RAG
von: Hu, Wentao, et al.
Veröffentlicht: (2025)
von: Hu, Wentao, et al.
Veröffentlicht: (2025)
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
von: Kulkarni, Atharva, et al.
Veröffentlicht: (2025)
von: Kulkarni, Atharva, et al.
Veröffentlicht: (2025)
A Scoping Review of Natural Language Processing in Addressing Medically Inaccurate Information: Errors, Misinformation, and Hallucination
von: Sun, Zhaoyi, et al.
Veröffentlicht: (2025)
von: Sun, Zhaoyi, et al.
Veröffentlicht: (2025)
Beyond Self-Consistency: Ensemble Reasoning Boosts Consistency and Accuracy of LLMs in Cancer Staging
von: Chang, Chia-Hsuan, et al.
Veröffentlicht: (2024)
von: Chang, Chia-Hsuan, et al.
Veröffentlicht: (2024)
The Two Sides of the Coin: Hallucination Generation and Detection with LLMs as Evaluators for LLMs
von: Bui, Anh Thu Maria, et al.
Veröffentlicht: (2024)
von: Bui, Anh Thu Maria, et al.
Veröffentlicht: (2024)
On A Scale From 1 to 5: Quantifying Hallucination in Faithfulness Evaluation
von: Jing, Xiaonan, et al.
Veröffentlicht: (2024)
von: Jing, Xiaonan, et al.
Veröffentlicht: (2024)
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
von: Tang, Liyan, et al.
Veröffentlicht: (2024)
von: Tang, Liyan, et al.
Veröffentlicht: (2024)
Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
Evaluating the Sensitivity of LLMs to Prior Context
von: Hankache, Robert, et al.
Veröffentlicht: (2025)
von: Hankache, Robert, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hallucinations and Truth: A Comprehensive Accuracy Evaluation of RAG, LoRA and DoRA
von: Baqar, Mohammad, et al.
Veröffentlicht: (2025) -
Beyond Facts: Evaluating Intent Hallucination in Large Language Models
von: Hao, Yijie, et al.
Veröffentlicht: (2025) -
Luna: An Evaluation Foundation Model to Catch Language Model Hallucinations with High Accuracy and Low Cost
von: Belyi, Masha, et al.
Veröffentlicht: (2024) -
Beyond Sentiment: A Multi-Agent Pipeline for Actionable Business Advice from Reviews
von: Bhandari, Kartikey Singh, et al.
Veröffentlicht: (2026) -
Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI
von: Li, Zihao, et al.
Veröffentlicht: (2025)