Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chatzoudis, Gerasimos, Polyzos, Konstantinos D., Li, Zhuowei, Gu, Difei, Moran, Gemma E., Wang, Hao, Metaxas, Dimitris N.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918447276883968
author Chatzoudis, Gerasimos
Polyzos, Konstantinos D.
Li, Zhuowei
Gu, Difei
Moran, Gemma E.
Wang, Hao
Metaxas, Dimitris N.
author_facet Chatzoudis, Gerasimos
Polyzos, Konstantinos D.
Li, Zhuowei
Gu, Difei
Moran, Gemma E.
Wang, Hao
Metaxas, Dimitris N.
contents Understanding the internal activations of Vision Transformers (ViTs) is critical for building interpretable and trustworthy models. While Sparse Autoencoders (SAEs) have been used to extract human-interpretable features, they operate on individual layers and fail to capture the cross-layer computational structure of Transformers, as well as the relative significance of each layer in forming the last-layer representation. Alternatively, we introduce the adoption of Cross-Layer Transcoders (CLTs) as reliable, sparse, and depth-aware proxy models for MLP blocks in ViTs. CLTs use an encoder-decoder scheme to reconstruct each post-MLP activation from learned sparse embeddings of preceding layers, yielding a linear decomposition that transforms the final representation of ViTs from an opaque embedding into an additive, layer-resolved construction that enables faithful attribution and process-level interpretability. We train CLTs on CLIP ViT-B/32 and ViT-B/16 across CIFAR-100, COCO, and ImageNet-100. We show that CLTs achieve high reconstruction fidelity with post-MLP activations while preserving and even improving, in some cases, CLIP zero-shot classification accuracy. In terms of interpretability, we show that the cross-layer contribution scores provide faithful attribution, revealing that the final representation is concentrated in a smaller set of dominant layer-wise terms whose removal degrades performance and whose retention largely preserves it. These results showcase the significance of adopting CLTs as an alternative interpretable proxy of ViTs in the vision domain.
format Preprint
id arxiv_https___arxiv_org_abs_2604_13304
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
Chatzoudis, Gerasimos
Polyzos, Konstantinos D.
Li, Zhuowei
Gu, Difei
Moran, Gemma E.
Wang, Hao
Metaxas, Dimitris N.
Computer Vision and Pattern Recognition
Artificial Intelligence
Understanding the internal activations of Vision Transformers (ViTs) is critical for building interpretable and trustworthy models. While Sparse Autoencoders (SAEs) have been used to extract human-interpretable features, they operate on individual layers and fail to capture the cross-layer computational structure of Transformers, as well as the relative significance of each layer in forming the last-layer representation. Alternatively, we introduce the adoption of Cross-Layer Transcoders (CLTs) as reliable, sparse, and depth-aware proxy models for MLP blocks in ViTs. CLTs use an encoder-decoder scheme to reconstruct each post-MLP activation from learned sparse embeddings of preceding layers, yielding a linear decomposition that transforms the final representation of ViTs from an opaque embedding into an additive, layer-resolved construction that enables faithful attribution and process-level interpretability. We train CLTs on CLIP ViT-B/32 and ViT-B/16 across CIFAR-100, COCO, and ImageNet-100. We show that CLTs achieve high reconstruction fidelity with post-MLP activations while preserving and even improving, in some cases, CLIP zero-shot classification accuracy. In terms of interpretability, we show that the cross-layer contribution scores provide faithful attribution, revealing that the final representation is concentrated in a smaller set of dominant layer-wise terms whose removal degrades performance and whose retention largely preserves it. These results showcase the significance of adopting CLTs as an alternative interpretable proxy of ViTs in the vision domain.
title Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2604.13304