EigenTrack: Spectral Activation Feature Tracking for Hallucination and Out-of-Distribution Detection in LLMs and VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ettori, Davide, Darabi, Nastaran, Tayebati, Sina, Krishnan, Ranganath, Subedar, Mahesh, Tickoo, Omesh, Trivedi, Amit Ranjan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917253378736128
author Ettori, Davide
Darabi, Nastaran
Tayebati, Sina
Krishnan, Ranganath
Subedar, Mahesh
Tickoo, Omesh
Trivedi, Amit Ranjan
author_facet Ettori, Davide
Darabi, Nastaran
Tayebati, Sina
Krishnan, Ranganath
Subedar, Mahesh
Tickoo, Omesh
Trivedi, Amit Ranjan
contents Large language models (LLMs) offer broad utility but remain prone to hallucination and out-of-distribution (OOD) errors. We propose EigenTrack, an interpretable real-time detector that uses the spectral geometry of hidden activations, a compact global signature of model dynamics. By streaming covariance-spectrum statistics such as entropy, eigenvalue gaps, and KL divergence from random baselines into a lightweight recurrent classifier, EigenTrack tracks temporal shifts in representation structure that signal hallucination and OOD drift before surface errors appear. Unlike black- and grey-box methods, it needs only a single forward pass without resampling. Unlike existing white-box detectors, it preserves temporal context, aggregates global signals, and offers interpretable accuracy-latency trade-offs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15735
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EigenTrack: Spectral Activation Feature Tracking for Hallucination and Out-of-Distribution Detection in LLMs and VLMs
Ettori, Davide
Darabi, Nastaran
Tayebati, Sina
Krishnan, Ranganath
Subedar, Mahesh
Tickoo, Omesh
Trivedi, Amit Ranjan
Machine Learning
Large language models (LLMs) offer broad utility but remain prone to hallucination and out-of-distribution (OOD) errors. We propose EigenTrack, an interpretable real-time detector that uses the spectral geometry of hidden activations, a compact global signature of model dynamics. By streaming covariance-spectrum statistics such as entropy, eigenvalue gaps, and KL divergence from random baselines into a lightweight recurrent classifier, EigenTrack tracks temporal shifts in representation structure that signal hallucination and OOD drift before surface errors appear. Unlike black- and grey-box methods, it needs only a single forward pass without resampling. Unlike existing white-box detectors, it preserves temporal context, aggregates global signals, and offers interpretable accuracy-latency trade-offs.
title EigenTrack: Spectral Activation Feature Tracking for Hallucination and Out-of-Distribution Detection in LLMs and VLMs
topic Machine Learning
url https://arxiv.org/abs/2509.15735