SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Abdaljalil, Samir, Pallucchini, Filippo, Seveso, Andrea, Kurban, Hasan, Mercorio, Fabio, Serpedin, Erchin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910859188502528
author Abdaljalil, Samir
Pallucchini, Filippo
Seveso, Andrea
Kurban, Hasan
Mercorio, Fabio
Serpedin, Erchin
author_facet Abdaljalil, Samir
Pallucchini, Filippo
Seveso, Andrea
Kurban, Hasan
Mercorio, Fabio
Serpedin, Erchin
contents Despite the state-of-the-art performance of Large Language Models (LLMs), these models often suffer from hallucinations, which can undermine their performance in critical applications. In this work, we propose SAFE, a novel method for detecting and mitigating hallucinations by leveraging Sparse Autoencoders (SAEs). While hallucination detection techniques and SAEs have been explored independently, their synergistic application in a comprehensive system, particularly for hallucination-aware query enrichment, has not been fully investigated. To validate the effectiveness of SAFE, we evaluate it on two models with available SAEs across three diverse cross-domain datasets designed to assess hallucination problems. Empirical results demonstrate that SAFE consistently improves query generation accuracy and mitigates hallucinations across all datasets, achieving accuracy improvements of up to 29.45%.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03032
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs
Abdaljalil, Samir
Pallucchini, Filippo
Seveso, Andrea
Kurban, Hasan
Mercorio, Fabio
Serpedin, Erchin
Computation and Language
Despite the state-of-the-art performance of Large Language Models (LLMs), these models often suffer from hallucinations, which can undermine their performance in critical applications. In this work, we propose SAFE, a novel method for detecting and mitigating hallucinations by leveraging Sparse Autoencoders (SAEs). While hallucination detection techniques and SAEs have been explored independently, their synergistic application in a comprehensive system, particularly for hallucination-aware query enrichment, has not been fully investigated. To validate the effectiveness of SAFE, we evaluate it on two models with available SAEs across three diverse cross-domain datasets designed to assess hallucination problems. Empirical results demonstrate that SAFE consistently improves query generation accuracy and mitigates hallucinations across all datasets, achieving accuracy improvements of up to 29.45%.
title SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs
topic Computation and Language
url https://arxiv.org/abs/2503.03032