Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Olson, Matthew Lyle, Hinck, Musashi, Ratzlaff, Neale, Li, Changbai, Howard, Phillip, Lal, Vasudev, Tseng, Shao-Yen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916751047917568
author Olson, Matthew Lyle
Hinck, Musashi
Ratzlaff, Neale
Li, Changbai
Howard, Phillip
Lal, Vasudev
Tseng, Shao-Yen
author_facet Olson, Matthew Lyle
Hinck, Musashi
Ratzlaff, Neale
Li, Changbai
Howard, Phillip
Lal, Vasudev
Tseng, Shao-Yen
contents The ImageNet hierarchy provides a structured taxonomy of object categories, offering a valuable lens through which to analyze the representations learned by deep vision models. In this work, we conduct a comprehensive analysis of how vision models encode the ImageNet hierarchy, leveraging Sparse Autoencoders (SAEs) to probe their internal representations. SAEs have been widely used as an explanation tool for large language models (LLMs), where they enable the discovery of semantically meaningful features. Here, we extend their use to vision models to investigate whether learned representations align with the ontological structure defined by the ImageNet taxonomy. Our results show that SAEs uncover hierarchical relationships in model activations, revealing an implicit encoding of taxonomic structure. We analyze the consistency of these representations across different layers of the popular vision foundation model DINOv2 and provide insights into how deep vision models internalize hierarchical category information by increasing information in the class token through each layer. Our study establishes a framework for systematic hierarchical analysis of vision model representations and highlights the potential of SAEs as a tool for probing semantic structure in deep networks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15970
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders
Olson, Matthew Lyle
Hinck, Musashi
Ratzlaff, Neale
Li, Changbai
Howard, Phillip
Lal, Vasudev
Tseng, Shao-Yen
Computer Vision and Pattern Recognition
Machine Learning
The ImageNet hierarchy provides a structured taxonomy of object categories, offering a valuable lens through which to analyze the representations learned by deep vision models. In this work, we conduct a comprehensive analysis of how vision models encode the ImageNet hierarchy, leveraging Sparse Autoencoders (SAEs) to probe their internal representations. SAEs have been widely used as an explanation tool for large language models (LLMs), where they enable the discovery of semantically meaningful features. Here, we extend their use to vision models to investigate whether learned representations align with the ontological structure defined by the ImageNet taxonomy. Our results show that SAEs uncover hierarchical relationships in model activations, revealing an implicit encoding of taxonomic structure. We analyze the consistency of these representations across different layers of the popular vision foundation model DINOv2 and provide insights into how deep vision models internalize hierarchical category information by increasing information in the class token through each layer. Our study establishes a framework for systematic hierarchical analysis of vision model representations and highlights the potential of SAEs as a tool for probing semantic structure in deep networks.
title Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2505.15970