Aligned Probing: Relating Toxic Behavior and Model Internals

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Waldis, Andreas, Gautam, Vagrant, Lauscher, Anne, Klakow, Dietrich, Gurevych, Iryna
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909803647860736
author Waldis, Andreas
Gautam, Vagrant
Lauscher, Anne
Klakow, Dietrich
Gurevych, Iryna
author_facet Waldis, Andreas
Gautam, Vagrant
Lauscher, Anne
Klakow, Dietrich
Gurevych, Iryna
contents We introduce aligned probing, a novel interpretability framework that aligns the behavior of language models (LMs), based on their outputs, and their internal representations (internals). Using this framework, we examine over 20 OLMo, Llama, and Mistral models, bridging behavioral and internal perspectives for toxicity for the first time. Our results show that LMs strongly encode information about the toxicity level of inputs and subsequent outputs, particularly in lower layers. Focusing on how unique LMs differ offers both correlative and causal evidence that they generate less toxic output when strongly encoding information about the input toxicity. We also highlight the heterogeneity of toxicity, as model behavior and internals vary across unique attributes such as Threat. Finally, four case studies analyzing detoxification, multi-prompt evaluations, model quantization, and pre-training dynamics underline the practical impact of aligned probing with further concrete insights. Our findings contribute to a more holistic understanding of LMs, both within and beyond the context of toxicity.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13390
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aligned Probing: Relating Toxic Behavior and Model Internals
Waldis, Andreas
Gautam, Vagrant
Lauscher, Anne
Klakow, Dietrich
Gurevych, Iryna
Computation and Language
We introduce aligned probing, a novel interpretability framework that aligns the behavior of language models (LMs), based on their outputs, and their internal representations (internals). Using this framework, we examine over 20 OLMo, Llama, and Mistral models, bridging behavioral and internal perspectives for toxicity for the first time. Our results show that LMs strongly encode information about the toxicity level of inputs and subsequent outputs, particularly in lower layers. Focusing on how unique LMs differ offers both correlative and causal evidence that they generate less toxic output when strongly encoding information about the input toxicity. We also highlight the heterogeneity of toxicity, as model behavior and internals vary across unique attributes such as Threat. Finally, four case studies analyzing detoxification, multi-prompt evaluations, model quantization, and pre-training dynamics underline the practical impact of aligned probing with further concrete insights. Our findings contribute to a more holistic understanding of LMs, both within and beyond the context of toxicity.
title Aligned Probing: Relating Toxic Behavior and Model Internals
topic Computation and Language
url https://arxiv.org/abs/2503.13390