Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908945236361216 |
|---|---|
| author | Hussain, Khizar Malin, Bradley A. Yin, Zhijun Rose, Susannah Leigh Kantarcioglu, Murat |
| author_facet | Hussain, Khizar Malin, Bradley A. Yin, Zhijun Rose, Susannah Leigh Kantarcioglu, Murat |
| contents | As LLM-powered chatbots are increasingly deployed in mental health services, detecting hallucinations and omissions has become critical for user safety. However, state-of-the-art LLM-as-a-judge methods often fail in high-risk healthcare
contexts, where subtle errors can have serious consequences. We show that leading LLM judges achieve only 52% accuracy on mental health counseling data, with some hallucination detection approaches exhibiting near-zero recall. We identify the root cause
as LLMs' inability to capture nuanced linguistic and therapeutic patterns recognized by domain experts. To address this, we propose a framework that integrates human expertise with LLMs to extract interpretable, domain-informed features across five
analytical dimensions: logical consistency, entity verification, factual accuracy, linguistic uncertainty, and professional appropriateness. Experiments on a public mental health dataset and a new human-annotated dataset show that traditional machine
learning models trained on these features achieve 0.717 F1 on our custom dataset and 0.849 F1 on a public benchmark for hallucination detection, with 0.59-0.64 F1 for omission detection across both datasets. Our results demonstrate that combining domain
expertise with automated methods yields more reliable and transparent evaluation than black-box LLM judging in high-stakes mental health applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_06216 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses Hussain, Khizar Malin, Bradley A. Yin, Zhijun Rose, Susannah Leigh Kantarcioglu, Murat Computation and Language Artificial Intelligence As LLM-powered chatbots are increasingly deployed in mental health services, detecting hallucinations and omissions has become critical for user safety. However, state-of-the-art LLM-as-a-judge methods often fail in high-risk healthcare contexts, where subtle errors can have serious consequences. We show that leading LLM judges achieve only 52% accuracy on mental health counseling data, with some hallucination detection approaches exhibiting near-zero recall. We identify the root cause as LLMs' inability to capture nuanced linguistic and therapeutic patterns recognized by domain experts. To address this, we propose a framework that integrates human expertise with LLMs to extract interpretable, domain-informed features across five analytical dimensions: logical consistency, entity verification, factual accuracy, linguistic uncertainty, and professional appropriateness. Experiments on a public mental health dataset and a new human-annotated dataset show that traditional machine learning models trained on these features achieve 0.717 F1 on our custom dataset and 0.849 F1 on a public benchmark for hallucination detection, with 0.59-0.64 F1 for omission detection across both datasets. Our results demonstrate that combining domain expertise with automated methods yields more reliable and transparent evaluation than black-box LLM judging in high-stakes mental health applications. |
| title | Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2604.06216 |