ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Zhongxiang, Zang, Xiaoxue, Zheng, Kai, Song, Yang, Xu, Jun, Zhang, Xiao, Yu, Weijie, Li, Han
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909462449618944
author Sun, Zhongxiang
Zang, Xiaoxue
Zheng, Kai
Song, Yang
Xu, Jun
Zhang, Xiao
Yu, Weijie
Song, Yang
Li, Han
author_facet Sun, Zhongxiang
Zang, Xiaoxue
Zheng, Kai
Song, Yang
Xu, Jun
Zhang, Xiao
Yu, Weijie
Song, Yang
Li, Han
contents Retrieval-Augmented Generation (RAG) models are designed to incorporate external knowledge, reducing hallucinations caused by insufficient parametric (internal) knowledge. However, even with accurate and relevant retrieved content, RAG models can still produce hallucinations by generating outputs that conflict with the retrieved information. Detecting such hallucinations requires disentangling how Large Language Models (LLMs) utilize external and parametric knowledge. Current detection methods often focus on one of these mechanisms or without decoupling their intertwined effects, making accurate detection difficult. In this paper, we investigate the internal mechanisms behind hallucinations in RAG scenarios. We discover hallucinations occur when the Knowledge FFNs in LLMs overemphasize parametric knowledge in the residual stream, while Copying Heads fail to effectively retain or integrate external knowledge from retrieved content. Based on these findings, we propose ReDeEP, a novel method that detects hallucinations by decoupling LLM's utilization of external context and parametric knowledge. Our experiments show that ReDeEP significantly improves RAG hallucination detection accuracy. Additionally, we introduce AARF, which mitigates hallucinations by modulating the contributions of Knowledge FFNs and Copying Heads.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11414
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability
Sun, Zhongxiang
Zang, Xiaoxue
Zheng, Kai
Song, Yang
Xu, Jun
Zhang, Xiao
Yu, Weijie
Song, Yang
Li, Han
Computation and Language
Retrieval-Augmented Generation (RAG) models are designed to incorporate external knowledge, reducing hallucinations caused by insufficient parametric (internal) knowledge. However, even with accurate and relevant retrieved content, RAG models can still produce hallucinations by generating outputs that conflict with the retrieved information. Detecting such hallucinations requires disentangling how Large Language Models (LLMs) utilize external and parametric knowledge. Current detection methods often focus on one of these mechanisms or without decoupling their intertwined effects, making accurate detection difficult. In this paper, we investigate the internal mechanisms behind hallucinations in RAG scenarios. We discover hallucinations occur when the Knowledge FFNs in LLMs overemphasize parametric knowledge in the residual stream, while Copying Heads fail to effectively retain or integrate external knowledge from retrieved content. Based on these findings, we propose ReDeEP, a novel method that detects hallucinations by decoupling LLM's utilization of external context and parametric knowledge. Our experiments show that ReDeEP significantly improves RAG hallucination detection accuracy. Additionally, we introduce AARF, which mitigates hallucinations by modulating the contributions of Knowledge FFNs and Copying Heads.
title ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability
topic Computation and Language
url https://arxiv.org/abs/2410.11414