Interpreting vision transformers via residual replacement model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Jinyeong, Kim, Junhyeok, Shim, Yumin, Kim, Joohyeok, Jung, Sunyoung, Hwang, Seong Jae
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914050506489856
author Kim, Jinyeong
Kim, Junhyeok
Shim, Yumin
Kim, Joohyeok
Jung, Sunyoung
Hwang, Seong Jae
author_facet Kim, Jinyeong
Kim, Junhyeok
Shim, Yumin
Kim, Joohyeok
Jung, Sunyoung
Hwang, Seong Jae
contents How do vision transformers (ViTs) represent and process the world? This paper addresses this long-standing question through the first systematic analysis of 6.6K features across all layers, extracted via sparse autoencoders, and by introducing the residual replacement model, which replaces ViT computations with interpretable features in the residual stream. Our analysis reveals not only a feature evolution from low-level patterns to high-level semantics, but also how ViTs encode curves and spatial positions through specialized feature types. The residual replacement model scalably produces a faithful yet parsimonious circuit for human-scale interpretability by significantly simplifying the original computations. As a result, this framework enables intuitive understanding of ViT mechanisms. Finally, we demonstrate the utility of our framework in debiasing spurious correlations.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17401
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Interpreting vision transformers via residual replacement model
Kim, Jinyeong
Kim, Junhyeok
Shim, Yumin
Kim, Joohyeok
Jung, Sunyoung
Hwang, Seong Jae
Computer Vision and Pattern Recognition
Artificial Intelligence
How do vision transformers (ViTs) represent and process the world? This paper addresses this long-standing question through the first systematic analysis of 6.6K features across all layers, extracted via sparse autoencoders, and by introducing the residual replacement model, which replaces ViT computations with interpretable features in the residual stream. Our analysis reveals not only a feature evolution from low-level patterns to high-level semantics, but also how ViTs encode curves and spatial positions through specialized feature types. The residual replacement model scalably produces a faithful yet parsimonious circuit for human-scale interpretability by significantly simplifying the original computations. As a result, this framework enables intuitive understanding of ViT mechanisms. Finally, we demonstrate the utility of our framework in debiasing spurious correlations.
title Interpreting vision transformers via residual replacement model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.17401