DePass: Unified Feature Attributing by Simple Decomposed Forward Pass

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hong, Xiangyu, Jiang, Che, Tian, Kai, Qi, Biqing, Sun, Youbang, Ding, Ning, Zhou, Bowen
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914111014567936
author Hong, Xiangyu
Jiang, Che
Tian, Kai
Qi, Biqing
Sun, Youbang
Ding, Ning
Zhou, Bowen
author_facet Hong, Xiangyu
Jiang, Che
Tian, Kai
Qi, Biqing
Sun, Youbang
Ding, Ning
Zhou, Bowen
contents Attributing the behavior of Transformer models to internal computations is a central challenge in mechanistic interpretability. We introduce DePass, a unified framework for feature attribution based on a single decomposed forward pass. DePass decomposes hidden states into customized additive components, then propagates them with attention scores and MLP's activations fixed. It achieves faithful, fine-grained attribution without requiring auxiliary training. We validate DePass across token-level, model component-level, and subspace-level attribution tasks, demonstrating its effectiveness and fidelity. Our experiments highlight its potential to attribute information flow between arbitrary components of a Transformer model. We hope DePass serves as a foundational tool for broader applications in interpretability.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18462
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DePass: Unified Feature Attributing by Simple Decomposed Forward Pass
Hong, Xiangyu
Jiang, Che
Tian, Kai
Qi, Biqing
Sun, Youbang
Ding, Ning
Zhou, Bowen
Computation and Language
Attributing the behavior of Transformer models to internal computations is a central challenge in mechanistic interpretability. We introduce DePass, a unified framework for feature attribution based on a single decomposed forward pass. DePass decomposes hidden states into customized additive components, then propagates them with attention scores and MLP's activations fixed. It achieves faithful, fine-grained attribution without requiring auxiliary training. We validate DePass across token-level, model component-level, and subspace-level attribution tasks, demonstrating its effectiveness and fidelity. Our experiments highlight its potential to attribute information flow between arbitrary components of a Transformer model. We hope DePass serves as a foundational tool for broader applications in interpretability.
title DePass: Unified Feature Attributing by Simple Decomposed Forward Pass
topic Computation and Language
url https://arxiv.org/abs/2510.18462