Saved in:
Bibliographic Details
Main Authors: Vachhani, Bhavik, Shrisvastava, Kush, Nema, Pranshu, Chiranthan, Sai
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.14829
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917413452251136
author Vachhani, Bhavik
Shrisvastava, Kush
Nema, Pranshu
Chiranthan, Sai
author_facet Vachhani, Bhavik
Shrisvastava, Kush
Nema, Pranshu
Chiranthan, Sai
contents Evaluating large language models (LLMs) for clinical documentation tasks such as SOAP note generation remains challenging. Unlike standard summarization, these tasks require clinical abstraction, normalization of colloquial language, and medically grounded inference. However, prevailing evaluation methods including automated metrics and LLM as judge frameworks rely on lexical faithfulness, often labeling any information not explicitly present in the transcript as hallucination. We show that such approaches systematically misclassify clinically valid outputs as errors, inflating hallucination rates and distorting model assessment. Our analysis reveals that many flagged hallucinations correspond to legitimate clinical transformations, including synonym mapping, abstraction of examination findings, diagnostic inference, and guideline consistent care planning. By aligning evaluation criteria with clinical reasoning through calibrated prompting and retrieval grounded in medical ontologies we observe a significant shift in outcomes. Under a lexical evaluation regime, the mean hallucination rate is 35%, heavily penalizing valid reasoning. With inference aware evaluation, this drops to 9%, with remaining cases reflecting genuine safety concerns. These findings suggest that current evaluation practices over penalize valid clinical reasoning and may measure artifacts of evaluation design rather than true errors, underscoring the need for clinically informed evaluation in high context domains like medicine.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14829
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Literal Summarization: Redefining Hallucination for Medical SOAP Note Evaluation
Vachhani, Bhavik
Shrisvastava, Kush
Nema, Pranshu
Chiranthan, Sai
Artificial Intelligence
Evaluating large language models (LLMs) for clinical documentation tasks such as SOAP note generation remains challenging. Unlike standard summarization, these tasks require clinical abstraction, normalization of colloquial language, and medically grounded inference. However, prevailing evaluation methods including automated metrics and LLM as judge frameworks rely on lexical faithfulness, often labeling any information not explicitly present in the transcript as hallucination. We show that such approaches systematically misclassify clinically valid outputs as errors, inflating hallucination rates and distorting model assessment. Our analysis reveals that many flagged hallucinations correspond to legitimate clinical transformations, including synonym mapping, abstraction of examination findings, diagnostic inference, and guideline consistent care planning. By aligning evaluation criteria with clinical reasoning through calibrated prompting and retrieval grounded in medical ontologies we observe a significant shift in outcomes. Under a lexical evaluation regime, the mean hallucination rate is 35%, heavily penalizing valid reasoning. With inference aware evaluation, this drops to 9%, with remaining cases reflecting genuine safety concerns. These findings suggest that current evaluation practices over penalize valid clinical reasoning and may measure artifacts of evaluation design rather than true errors, underscoring the need for clinically informed evaluation in high context domains like medicine.
title Beyond Literal Summarization: Redefining Hallucination for Medical SOAP Note Evaluation
topic Artificial Intelligence
url https://arxiv.org/abs/2604.14829