Hallucination Detection in LLMs with Topological Divergence on Attention Graphs
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866913114503512064 |
|---|---|
| author | Bazarova, Alexandra Volodichev, Andrei Yugay, Aleksandr Shulga, Andrey Ermilova, Alina Polev, Konstantin Belikova, Julia Parchiev, Rauf Simakov, Dmitry Savchenko, Maxim Savchenko, Andrey Barannikov, Serguei Zaytsev, Alexey |
| author_facet | Bazarova, Alexandra Volodichev, Andrei Yugay, Aleksandr Shulga, Andrey Ermilova, Alina Polev, Konstantin Belikova, Julia Parchiev, Rauf Simakov, Dmitry Savchenko, Maxim Savchenko, Andrey Barannikov, Serguei Zaytsev, Alexey |
| contents | Hallucination, i.e., generating factually incorrect content, remains a critical challenge for large language models (LLMs). We introduce TOHA, a TOpology-based HAllucination detector in the RAG setting, which leverages a topological divergence metric to quantify the structural properties of graphs induced by attention matrices. Examining the topological divergence between prompt and response subgraphs reveals consistent patterns: higher divergence values in specific attention heads correlate with hallucinated outputs, independent of the dataset. Extensive experiments - including evaluation on question answering and summarization tasks - show that our approach achieves state-of-the-art or competitive results on several benchmarks while requiring minimal annotated data and computational resources. Our findings suggest that analyzing the topological structure of attention matrices can serve as an efficient and robust indicator of factual reliability in LLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_10063 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Hallucination Detection in LLMs with Topological Divergence on Attention Graphs Bazarova, Alexandra Volodichev, Andrei Yugay, Aleksandr Shulga, Andrey Ermilova, Alina Polev, Konstantin Belikova, Julia Parchiev, Rauf Simakov, Dmitry Savchenko, Maxim Savchenko, Andrey Barannikov, Serguei Zaytsev, Alexey Computation and Language Artificial Intelligence Hallucination, i.e., generating factually incorrect content, remains a critical challenge for large language models (LLMs). We introduce TOHA, a TOpology-based HAllucination detector in the RAG setting, which leverages a topological divergence metric to quantify the structural properties of graphs induced by attention matrices. Examining the topological divergence between prompt and response subgraphs reveals consistent patterns: higher divergence values in specific attention heads correlate with hallucinated outputs, independent of the dataset. Extensive experiments - including evaluation on question answering and summarization tasks - show that our approach achieves state-of-the-art or competitive results on several benchmarks while requiring minimal annotated data and computational resources. Our findings suggest that analyzing the topological structure of attention matrices can serve as an efficient and robust indicator of factual reliability in LLMs. |
| title | Hallucination Detection in LLMs with Topological Divergence on Attention Graphs |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2504.10063 |