Hallucination Detection in LLMs with Topological Divergence on Attention Graphs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Bazarova, Alexandra, Volodichev, Andrei, Yugay, Aleksandr, Shulga, Andrey, Ermilova, Alina, Polev, Konstantin, Belikova, Julia, Parchiev, Rauf, Simakov, Dmitry, Savchenko, Maxim, Savchenko, Andrey, Barannikov, Serguei, Zaytsev, Alexey
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913114503512064
author Bazarova, Alexandra
Volodichev, Andrei
Yugay, Aleksandr
Shulga, Andrey
Ermilova, Alina
Polev, Konstantin
Belikova, Julia
Parchiev, Rauf
Simakov, Dmitry
Savchenko, Maxim
Savchenko, Andrey
Barannikov, Serguei
Zaytsev, Alexey
author_facet Bazarova, Alexandra
Volodichev, Andrei
Yugay, Aleksandr
Shulga, Andrey
Ermilova, Alina
Polev, Konstantin
Belikova, Julia
Parchiev, Rauf
Simakov, Dmitry
Savchenko, Maxim
Savchenko, Andrey
Barannikov, Serguei
Zaytsev, Alexey
contents Hallucination, i.e., generating factually incorrect content, remains a critical challenge for large language models (LLMs). We introduce TOHA, a TOpology-based HAllucination detector in the RAG setting, which leverages a topological divergence metric to quantify the structural properties of graphs induced by attention matrices. Examining the topological divergence between prompt and response subgraphs reveals consistent patterns: higher divergence values in specific attention heads correlate with hallucinated outputs, independent of the dataset. Extensive experiments - including evaluation on question answering and summarization tasks - show that our approach achieves state-of-the-art or competitive results on several benchmarks while requiring minimal annotated data and computational resources. Our findings suggest that analyzing the topological structure of attention matrices can serve as an efficient and robust indicator of factual reliability in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10063
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hallucination Detection in LLMs with Topological Divergence on Attention Graphs
Bazarova, Alexandra
Volodichev, Andrei
Yugay, Aleksandr
Shulga, Andrey
Ermilova, Alina
Polev, Konstantin
Belikova, Julia
Parchiev, Rauf
Simakov, Dmitry
Savchenko, Maxim
Savchenko, Andrey
Barannikov, Serguei
Zaytsev, Alexey
Computation and Language
Artificial Intelligence
Hallucination, i.e., generating factually incorrect content, remains a critical challenge for large language models (LLMs). We introduce TOHA, a TOpology-based HAllucination detector in the RAG setting, which leverages a topological divergence metric to quantify the structural properties of graphs induced by attention matrices. Examining the topological divergence between prompt and response subgraphs reveals consistent patterns: higher divergence values in specific attention heads correlate with hallucinated outputs, independent of the dataset. Extensive experiments - including evaluation on question answering and summarization tasks - show that our approach achieves state-of-the-art or competitive results on several benchmarks while requiring minimal annotated data and computational resources. Our findings suggest that analyzing the topological structure of attention matrices can serve as an efficient and robust indicator of factual reliability in LLMs.
title Hallucination Detection in LLMs with Topological Divergence on Attention Graphs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.10063