Interpreting vision transformers via residual replacement model
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914050506489856 |
|---|---|
| author | Kim, Jinyeong Kim, Junhyeok Shim, Yumin Kim, Joohyeok Jung, Sunyoung Hwang, Seong Jae |
| author_facet | Kim, Jinyeong Kim, Junhyeok Shim, Yumin Kim, Joohyeok Jung, Sunyoung Hwang, Seong Jae |
| contents | How do vision transformers (ViTs) represent and process the world? This paper addresses this long-standing question through the first systematic analysis of 6.6K features across all layers, extracted via sparse autoencoders, and by introducing the residual replacement model, which replaces ViT computations with interpretable features in the residual stream. Our analysis reveals not only a feature evolution from low-level patterns to high-level semantics, but also how ViTs encode curves and spatial positions through specialized feature types. The residual replacement model scalably produces a faithful yet parsimonious circuit for human-scale interpretability by significantly simplifying the original computations. As a result, this framework enables intuitive understanding of ViT mechanisms. Finally, we demonstrate the utility of our framework in debiasing spurious correlations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_17401 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Interpreting vision transformers via residual replacement model Kim, Jinyeong Kim, Junhyeok Shim, Yumin Kim, Joohyeok Jung, Sunyoung Hwang, Seong Jae Computer Vision and Pattern Recognition Artificial Intelligence How do vision transformers (ViTs) represent and process the world? This paper addresses this long-standing question through the first systematic analysis of 6.6K features across all layers, extracted via sparse autoencoders, and by introducing the residual replacement model, which replaces ViT computations with interpretable features in the residual stream. Our analysis reveals not only a feature evolution from low-level patterns to high-level semantics, but also how ViTs encode curves and spatial positions through specialized feature types. The residual replacement model scalably produces a faithful yet parsimonious circuit for human-scale interpretability by significantly simplifying the original computations. As a result, this framework enables intuitive understanding of ViT mechanisms. Finally, we demonstrate the utility of our framework in debiasing spurious correlations. |
| title | Interpreting vision transformers via residual replacement model |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2509.17401 |