LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Jiarui, Jain, Jivitesh, Diab, Mona, Subramani, Nishant
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918154588913664
author Liu, Jiarui
Jain, Jivitesh
Diab, Mona
Subramani, Nishant
author_facet Liu, Jiarui
Jain, Jivitesh
Diab, Mona
Subramani, Nishant
contents Although large language models (LLMs) have tremendous utility, trustworthiness is still a chief concern: models often generate incorrect information with high confidence. While contextual information can help guide generation, identifying when a query would benefit from retrieved context and assessing the effectiveness of that context remains challenging. In this work, we operationalize interpretability methods to ascertain whether we can predict the correctness of model outputs from the model's activations alone. We also explore whether model internals contain signals about the efficacy of external context. We consider correct, incorrect, and irrelevant context and introduce metrics to distinguish amongst them. Experiments on six different models reveal that a simple classifier trained on intermediate layer activations of the first output token can predict output correctness with about 75% accuracy, enabling early auditing. Our model-internals-based metric significantly outperforms prompting baselines at distinguishing between correct and incorrect context, guarding against inaccuracies introduced by polluted context. These findings offer a lens to better understand the underlying decision-making processes of LLMs. Our code is publicly available at https://github.com/jiarui-liu/LLM-Microscope
format Preprint
id arxiv_https___arxiv_org_abs_2510_04013
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
Liu, Jiarui
Jain, Jivitesh
Diab, Mona
Subramani, Nishant
Computation and Language
Although large language models (LLMs) have tremendous utility, trustworthiness is still a chief concern: models often generate incorrect information with high confidence. While contextual information can help guide generation, identifying when a query would benefit from retrieved context and assessing the effectiveness of that context remains challenging. In this work, we operationalize interpretability methods to ascertain whether we can predict the correctness of model outputs from the model's activations alone. We also explore whether model internals contain signals about the efficacy of external context. We consider correct, incorrect, and irrelevant context and introduce metrics to distinguish amongst them. Experiments on six different models reveal that a simple classifier trained on intermediate layer activations of the first output token can predict output correctness with about 75% accuracy, enabling early auditing. Our model-internals-based metric significantly outperforms prompting baselines at distinguishing between correct and incorrect context, guarding against inaccuracies introduced by polluted context. These findings offer a lens to better understand the underlying decision-making processes of LLMs. Our code is publicly available at https://github.com/jiarui-liu/LLM-Microscope
title LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
topic Computation and Language
url https://arxiv.org/abs/2510.04013