From Colors to Classes: Emergence of Concepts in Vision Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dorszewski, Teresa, Tětková, Lenka, Jenssen, Robert, Hansen, Lars Kai, Wickstrøm, Kristoffer Knutsen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916667886403584
author Dorszewski, Teresa
Tětková, Lenka
Jenssen, Robert
Hansen, Lars Kai
Wickstrøm, Kristoffer Knutsen
author_facet Dorszewski, Teresa
Tětková, Lenka
Jenssen, Robert
Hansen, Lars Kai
Wickstrøm, Kristoffer Knutsen
contents Vision Transformers (ViTs) are increasingly utilized in various computer vision tasks due to their powerful representation capabilities. However, it remains understudied how ViTs process information layer by layer. Numerous studies have shown that convolutional neural networks (CNNs) extract features of increasing complexity throughout their layers, which is crucial for tasks like domain adaptation and transfer learning. ViTs, lacking the same inductive biases as CNNs, can potentially learn global dependencies from the first layers due to their attention mechanisms. Given the increasing importance of ViTs in computer vision, there is a need to improve the layer-wise understanding of ViTs. In this work, we present a novel, layer-wise analysis of concepts encoded in state-of-the-art ViTs using neuron labeling. Our findings reveal that ViTs encode concepts with increasing complexity throughout the network. Early layers primarily encode basic features such as colors and textures, while later layers represent more specific classes, including objects and animals. As the complexity of encoded concepts increases, the number of concepts represented in each layer also rises, reflecting a more diverse and specific set of features. Additionally, different pretraining strategies influence the quantity and category of encoded concepts, with finetuning to specific downstream tasks generally reducing the number of encoded concepts and shifting the concepts to more relevant categories.
format Preprint
id arxiv_https___arxiv_org_abs_2503_24071
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Colors to Classes: Emergence of Concepts in Vision Transformers
Dorszewski, Teresa
Tětková, Lenka
Jenssen, Robert
Hansen, Lars Kai
Wickstrøm, Kristoffer Knutsen
Computer Vision and Pattern Recognition
Machine Learning
Vision Transformers (ViTs) are increasingly utilized in various computer vision tasks due to their powerful representation capabilities. However, it remains understudied how ViTs process information layer by layer. Numerous studies have shown that convolutional neural networks (CNNs) extract features of increasing complexity throughout their layers, which is crucial for tasks like domain adaptation and transfer learning. ViTs, lacking the same inductive biases as CNNs, can potentially learn global dependencies from the first layers due to their attention mechanisms. Given the increasing importance of ViTs in computer vision, there is a need to improve the layer-wise understanding of ViTs. In this work, we present a novel, layer-wise analysis of concepts encoded in state-of-the-art ViTs using neuron labeling. Our findings reveal that ViTs encode concepts with increasing complexity throughout the network. Early layers primarily encode basic features such as colors and textures, while later layers represent more specific classes, including objects and animals. As the complexity of encoded concepts increases, the number of concepts represented in each layer also rises, reflecting a more diverse and specific set of features. Additionally, different pretraining strategies influence the quantity and category of encoded concepts, with finetuning to specific downstream tasks generally reducing the number of encoded concepts and shifting the concepts to more relevant categories.
title From Colors to Classes: Emergence of Concepts in Vision Transformers
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2503.24071