Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yan, Xinyuan, Liu, Shusen, Thopalli, Kowshik, Wang, Bei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911255598465024
author Yan, Xinyuan
Liu, Shusen
Thopalli, Kowshik
Wang, Bei
author_facet Yan, Xinyuan
Liu, Shusen
Thopalli, Kowshik
Wang, Bei
contents Sparse autoencoders (SAEs) have emerged as a powerful tool for uncovering interpretable features in large language models (LLMs) through the sparse directions they learn. However, the sheer number of extracted directions makes comprehensive exploration intractable. While conventional embedding techniques such as UMAP can reveal global structure, they suffer from limitations including high-dimensional compression artifacts, overplotting, and misleading neighborhood distortions. In this work, we propose a focused exploration framework that prioritizes curated concepts and their corresponding SAE features over attempts to visualize all available features simultaneously. We present an interactive visualization system that combines topology-based visual encoding with dimensionality reduction to faithfully represent both local and global relationships among selected features. This hybrid approach enables users to investigate SAE behavior through targeted, interpretable subsets, facilitating deeper and more nuanced analysis of concept representation in latent space.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06048
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
Yan, Xinyuan
Liu, Shusen
Thopalli, Kowshik
Wang, Bei
Computation and Language
Machine Learning
Sparse autoencoders (SAEs) have emerged as a powerful tool for uncovering interpretable features in large language models (LLMs) through the sparse directions they learn. However, the sheer number of extracted directions makes comprehensive exploration intractable. While conventional embedding techniques such as UMAP can reveal global structure, they suffer from limitations including high-dimensional compression artifacts, overplotting, and misleading neighborhood distortions. In this work, we propose a focused exploration framework that prioritizes curated concepts and their corresponding SAE features over attempts to visualize all available features simultaneously. We present an interactive visualization system that combines topology-based visual encoding with dimensionality reduction to faithfully represent both local and global relationships among selected features. This hybrid approach enables users to investigate SAE behavior through targeted, interpretable subsets, facilitating deeper and more nuanced analysis of concept representation in latent space.
title Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2511.06048