ReXTrust: A Model for Fine-Grained Hallucination Detection in AI-Generated Radiology Reports
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909470672551936 |
|---|---|
| author | Hardy, Romain Kim, Sung Eun Ro, Du Hyun Rajpurkar, Pranav |
| author_facet | Hardy, Romain Kim, Sung Eun Ro, Du Hyun Rajpurkar, Pranav |
| contents | The increasing adoption of AI-generated radiology reports necessitates robust methods for detecting hallucinations--false or unfounded statements that could impact patient care. We present ReXTrust, a novel framework for fine-grained hallucination detection in AI-generated radiology reports. Our approach leverages sequences of hidden states from large vision-language models to produce finding-level hallucination risk scores. We evaluate ReXTrust on a subset of the MIMIC-CXR dataset and demonstrate superior performance compared to existing approaches, achieving an AUROC of 0.8751 across all findings and 0.8963 on clinically significant findings. Our results show that white-box approaches leveraging model hidden states can provide reliable hallucination detection for medical AI systems, potentially improving the safety and reliability of automated radiology reporting. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_15264 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ReXTrust: A Model for Fine-Grained Hallucination Detection in AI-Generated Radiology Reports Hardy, Romain Kim, Sung Eun Ro, Du Hyun Rajpurkar, Pranav Computation and Language Artificial Intelligence The increasing adoption of AI-generated radiology reports necessitates robust methods for detecting hallucinations--false or unfounded statements that could impact patient care. We present ReXTrust, a novel framework for fine-grained hallucination detection in AI-generated radiology reports. Our approach leverages sequences of hidden states from large vision-language models to produce finding-level hallucination risk scores. We evaluate ReXTrust on a subset of the MIMIC-CXR dataset and demonstrate superior performance compared to existing approaches, achieving an AUROC of 0.8751 across all findings and 0.8963 on clinically significant findings. Our results show that white-box approaches leveraging model hidden states can provide reliable hallucination detection for medical AI systems, potentially improving the safety and reliability of automated radiology reporting. |
| title | ReXTrust: A Model for Fine-Grained Hallucination Detection in AI-Generated Radiology Reports |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2412.15264 |