Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses
Fuente:
arXiv
Saved in:
| Main Authors: | Hussain, Khizar, Malin, Bradley A., Yin, Zhijun, Rose, Susannah Leigh, Kantarcioglu, Murat |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PARALLAX: Separating Genuine Hallucination Detection from Benchmark Construction Artifacts
by: Hussain, Khizar, et al.
Published: (2026)
by: Hussain, Khizar, et al.
Published: (2026)
Disentangling Prompt Element Level Risk Factors for Hallucinations and Omissions in Mental Health LLM Responses
by: Ni, Congning, et al.
Published: (2026)
by: Ni, Congning, et al.
Published: (2026)
MHGraphBench: Knowledge Graph-Grounded Benchmarking of Mental Health Knowledge in Large Language Models
by: Liu, Weixin, et al.
Published: (2026)
by: Liu, Weixin, et al.
Published: (2026)
VERI-DPO: Evidence-Aware Alignment for Clinical Summarization via Claim Verification and Direct Preference Optimization
by: Liu, Weixin, et al.
Published: (2026)
by: Liu, Weixin, et al.
Published: (2026)
Enhancing Knowledge Graph Construction: Evaluating with Emphasis on Hallucination, Omission, and Graph Similarity Metrics
by: Ghanem, Hussam, et al.
Published: (2025)
by: Ghanem, Hussam, et al.
Published: (2025)
Integrating Large Language Models with Human Expertise for Disease Detection in Electronic Health Records
by: Pan, Jie, et al.
Published: (2025)
by: Pan, Jie, et al.
Published: (2025)
Mitigating Hallucinations Using Ensemble of Knowledge Graph and Vector Store in Large Language Models to Enhance Mental Health Support
by: Muqtadir, Abdul, et al.
Published: (2024)
by: Muqtadir, Abdul, et al.
Published: (2024)
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
by: Guo, Kevin H., et al.
Published: (2026)
by: Guo, Kevin H., et al.
Published: (2026)
Stop Listening to Me! How Multi-turn Conversations Can Degrade LLM Reliability
by: Guo, Kevin H., et al.
Published: (2026)
by: Guo, Kevin H., et al.
Published: (2026)
DialogueForge: LLM Simulation of Human-Chatbot Dialogue
by: Zhu, Ruizhe, et al.
Published: (2025)
by: Zhu, Ruizhe, et al.
Published: (2025)
SHROOM-INDElab at SemEval-2024 Task 6: Zero- and Few-Shot LLM-Based Classification for Hallucination Detection
by: Allen, Bradley P., et al.
Published: (2024)
by: Allen, Bradley P., et al.
Published: (2024)
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools
by: Park, Jung In, et al.
Published: (2024)
by: Park, Jung In, et al.
Published: (2024)
Extrinsically-Focused Evaluation of Omissions in Medical Summarization
by: Schumacher, Elliot, et al.
Published: (2023)
by: Schumacher, Elliot, et al.
Published: (2023)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
by: Atasoy, I. F., et al.
Published: (2026)
by: Atasoy, I. F., et al.
Published: (2026)
From Physician Expertise to Clinical Agents: Preserving, Standardizing, and Scaling Physicians' Medical Expertise with Lightweight LLM
by: Luo, Chanyong, et al.
Published: (2026)
by: Luo, Chanyong, et al.
Published: (2026)
Analyzing LLM Reasoning to Uncover Mental Health Stigma
by: Sankar, Sreehari, et al.
Published: (2026)
by: Sankar, Sreehari, et al.
Published: (2026)
Bypassing AI Control Protocols via Agent-as-a-Proxy Attacks
by: Isbarov, Jafar, et al.
Published: (2026)
by: Isbarov, Jafar, et al.
Published: (2026)
Steer LLM Latents for Hallucination Detection
by: Park, Seongheon, et al.
Published: (2025)
by: Park, Seongheon, et al.
Published: (2025)
MedAide: Leveraging Large Language Models for On-Premise Medical Assistance on Edge Devices
by: Basit, Abdul, et al.
Published: (2024)
by: Basit, Abdul, et al.
Published: (2024)
Large Language Model for Mental Health: A Systematic Review
by: Guo, Zhijun, et al.
Published: (2024)
by: Guo, Zhijun, et al.
Published: (2024)
CLEAR: Revealing How Noise and Ambiguity Degrade Reliability in LLMs for Medicine
by: Guo, Kevin H., et al.
Published: (2026)
by: Guo, Kevin H., et al.
Published: (2026)
Citation-Enhanced Generation for LLM-based Chatbots
by: Li, Weitao, et al.
Published: (2024)
by: Li, Weitao, et al.
Published: (2024)
The Knowledge-Behaviour Disconnect in LLM-based Chatbots
by: Broersen, Jan
Published: (2025)
by: Broersen, Jan
Published: (2025)
Hallucination Detection and Hallucination Mitigation: An Investigation
by: Luo, Junliang, et al.
Published: (2024)
by: Luo, Junliang, et al.
Published: (2024)
Budget-Aware Routing for Long Clinical Text
by: Qureshi, Khizar, et al.
Published: (2026)
by: Qureshi, Khizar, et al.
Published: (2026)
A Complete Survey on LLM-based AI Chatbots
by: Dam, Sumit Kumar, et al.
Published: (2024)
by: Dam, Sumit Kumar, et al.
Published: (2024)
Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection
by: Gao, Weizhi, et al.
Published: (2025)
by: Gao, Weizhi, et al.
Published: (2025)
TherapyProbe: Generating Design Knowledge for Relational Safety in Mental Health Chatbots Through Adversarial Simulation
by: Chandra, Joydeep, et al.
Published: (2026)
by: Chandra, Joydeep, et al.
Published: (2026)
Cognitive-Mental-LLM: Evaluating Reasoning in Large Language Models for Mental Health Prediction via Online Text
by: Patil, Avinash, et al.
Published: (2025)
by: Patil, Avinash, et al.
Published: (2025)
Promoting the Responsible Development of Speech Datasets for Mental Health and Neurological Disorders Research
by: Mancini, Eleonora, et al.
Published: (2024)
by: Mancini, Eleonora, et al.
Published: (2024)
Human Decision-Making with Persuasive and Narrative LLM Explanations
by: Marusich, Laura R., et al.
Published: (2026)
by: Marusich, Laura R., et al.
Published: (2026)
HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses
by: Lee, Sujeong, et al.
Published: (2025)
by: Lee, Sujeong, et al.
Published: (2025)
Is LLMs Hallucination Usable? LLM-based Negative Reasoning for Fake News Detection
by: Zhang, Chaowei, et al.
Published: (2025)
by: Zhang, Chaowei, et al.
Published: (2025)
PsychAdapter: Adapting LLM Transformers to Reflect Traits, Personality and Mental Health
by: Vu, Huy, et al.
Published: (2024)
by: Vu, Huy, et al.
Published: (2024)
Enhancing Hallucination Detection through Perturbation-Based Synthetic Data Generation in System Responses
by: Zhang, Dongxu, et al.
Published: (2024)
by: Zhang, Dongxu, et al.
Published: (2024)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
by: Chiang, Wei-Lin, et al.
Published: (2024)
by: Chiang, Wei-Lin, et al.
Published: (2024)
Faithfulness metric fusion: Improving the evaluation of LLM trustworthiness across domains
by: Malin, Ben, et al.
Published: (2025)
by: Malin, Ben, et al.
Published: (2025)
HalluLens: LLM Hallucination Benchmark
by: Bang, Yejin, et al.
Published: (2025)
by: Bang, Yejin, et al.
Published: (2025)
The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues
by: Cheng, Youyou, et al.
Published: (2026)
by: Cheng, Youyou, et al.
Published: (2026)
ToBlend: Token-Level Blending With an Ensemble of LLMs to Attack AI-Generated Text Detection
by: Huang, Fan, et al.
Published: (2024)
by: Huang, Fan, et al.
Published: (2024)
Similar Items
-
PARALLAX: Separating Genuine Hallucination Detection from Benchmark Construction Artifacts
by: Hussain, Khizar, et al.
Published: (2026) -
Disentangling Prompt Element Level Risk Factors for Hallucinations and Omissions in Mental Health LLM Responses
by: Ni, Congning, et al.
Published: (2026) -
MHGraphBench: Knowledge Graph-Grounded Benchmarking of Mental Health Knowledge in Large Language Models
by: Liu, Weixin, et al.
Published: (2026) -
VERI-DPO: Evidence-Aware Alignment for Clinical Summarization via Claim Verification and Direct Preference Optimization
by: Liu, Weixin, et al.
Published: (2026) -
Enhancing Knowledge Graph Construction: Evaluating with Emphasis on Hallucination, Omission, and Graph Similarity Metrics
by: Ghanem, Hussam, et al.
Published: (2025)