Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Yu-Ting, Chang, Fu-Chieh, Shu, Yu-En, Shih, Hui-Ying, Wu, Pei-Yuan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915791876653056
author Lee, Yu-Ting
Chang, Fu-Chieh
Shu, Yu-En
Shih, Hui-Ying
Wu, Pei-Yuan
author_facet Lee, Yu-Ting
Chang, Fu-Chieh
Shu, Yu-En
Shih, Hui-Ying
Wu, Pei-Yuan
contents Intrinsic self-correction refers to the phenomenon where a language model refines its own outputs purely through prompting, without external feedback or parameter updates. While this approach improves performance across diverse tasks, its mechanism remains unclear. We show that intrinsic self-correction functions by steering hidden representations along interpretable latent directions, as evidenced by both alignment analysis and activation interventions. To achieve this, we analyze intrinsic self-correction via the representation shift induced by prompting. In parallel, we construct interpretable latent directions with contrastive pairs and verify the causal effect of these directions via activation addition. Evaluating six open-source LLMs, our results demonstrate that prompt-induced representation shifts in text detoxification and text toxification consistently align with latent directions constructed from contrastive pairs. In detoxification, the shifts align with the non-toxic direction; in toxification, they align with the toxic direction. These findings suggest that representation steering is the mechanistic driver of intrinsic self-correction. Our analysis highlights that understanding model internals offers a direct route to analyzing the mechanisms of prompt-driven LLM behaviors.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11924
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
Lee, Yu-Ting
Chang, Fu-Chieh
Shu, Yu-En
Shih, Hui-Ying
Wu, Pei-Yuan
Computation and Language
Artificial Intelligence
Machine Learning
Intrinsic self-correction refers to the phenomenon where a language model refines its own outputs purely through prompting, without external feedback or parameter updates. While this approach improves performance across diverse tasks, its mechanism remains unclear. We show that intrinsic self-correction functions by steering hidden representations along interpretable latent directions, as evidenced by both alignment analysis and activation interventions. To achieve this, we analyze intrinsic self-correction via the representation shift induced by prompting. In parallel, we construct interpretable latent directions with contrastive pairs and verify the causal effect of these directions via activation addition. Evaluating six open-source LLMs, our results demonstrate that prompt-induced representation shifts in text detoxification and text toxification consistently align with latent directions constructed from contrastive pairs. In detoxification, the shifts align with the non-toxic direction; in toxification, they align with the toxic direction. These findings suggest that representation steering is the mechanistic driver of intrinsic self-correction. Our analysis highlights that understanding model internals offers a direct route to analyzing the mechanisms of prompt-driven LLM behaviors.
title Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.11924