DePass: Unified Feature Attributing by Simple Decomposed Forward Pass
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914111014567936 |
|---|---|
| author | Hong, Xiangyu Jiang, Che Tian, Kai Qi, Biqing Sun, Youbang Ding, Ning Zhou, Bowen |
| author_facet | Hong, Xiangyu Jiang, Che Tian, Kai Qi, Biqing Sun, Youbang Ding, Ning Zhou, Bowen |
| contents | Attributing the behavior of Transformer models to internal computations is a central challenge in mechanistic interpretability. We introduce DePass, a unified framework for feature attribution based on a single decomposed forward pass. DePass decomposes hidden states into customized additive components, then propagates them with attention scores and MLP's activations fixed. It achieves faithful, fine-grained attribution without requiring auxiliary training. We validate DePass across token-level, model component-level, and subspace-level attribution tasks, demonstrating its effectiveness and fidelity. Our experiments highlight its potential to attribute information flow between arbitrary components of a Transformer model. We hope DePass serves as a foundational tool for broader applications in interpretability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_18462 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DePass: Unified Feature Attributing by Simple Decomposed Forward Pass Hong, Xiangyu Jiang, Che Tian, Kai Qi, Biqing Sun, Youbang Ding, Ning Zhou, Bowen Computation and Language Attributing the behavior of Transformer models to internal computations is a central challenge in mechanistic interpretability. We introduce DePass, a unified framework for feature attribution based on a single decomposed forward pass. DePass decomposes hidden states into customized additive components, then propagates them with attention scores and MLP's activations fixed. It achieves faithful, fine-grained attribution without requiring auxiliary training. We validate DePass across token-level, model component-level, and subspace-level attribution tasks, demonstrating its effectiveness and fidelity. Our experiments highlight its potential to attribute information flow between arbitrary components of a Transformer model. We hope DePass serves as a foundational tool for broader applications in interpretability. |
| title | DePass: Unified Feature Attributing by Simple Decomposed Forward Pass |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.18462 |