The Deleuzian Representation Hypothesis

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cornet, Clément, Besançon, Romaric, Borgne, Hervé Le
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910019224600576
author Cornet, Clément
Besançon, Romaric
Borgne, Hervé Le
author_facet Cornet, Clément
Besançon, Romaric
Borgne, Hervé Le
contents We propose an alternative to sparse autoencoders (SAEs) as a simple and effective unsupervised method for extracting interpretable concepts from neural networks. The core idea is to cluster differences in activations, which we formally justify within a discriminant analysis framework. To enhance the diversity of extracted concepts, we refine the approach by weighting the clustering using the skewness of activations. The method aligns with Deleuze's modern view of concepts as differences. We evaluate the approach across five models and three modalities (vision, language, and audio), measuring concept quality, diversity, and consistency. Our results show that the proposed method achieves concept quality surpassing prior unsupervised SAE variants while approaching supervised baselines, and that the extracted concepts enable steering of a model's inner representations, demonstrating their causal influence on downstream behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19734
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Deleuzian Representation Hypothesis
Cornet, Clément
Besançon, Romaric
Borgne, Hervé Le
Machine Learning
We propose an alternative to sparse autoencoders (SAEs) as a simple and effective unsupervised method for extracting interpretable concepts from neural networks. The core idea is to cluster differences in activations, which we formally justify within a discriminant analysis framework. To enhance the diversity of extracted concepts, we refine the approach by weighting the clustering using the skewness of activations. The method aligns with Deleuze's modern view of concepts as differences. We evaluate the approach across five models and three modalities (vision, language, and audio), measuring concept quality, diversity, and consistency. Our results show that the proposed method achieves concept quality surpassing prior unsupervised SAE variants while approaching supervised baselines, and that the extracted concepts enable steering of a model's inner representations, demonstrating their causal influence on downstream behavior.
title The Deleuzian Representation Hypothesis
topic Machine Learning
url https://arxiv.org/abs/2512.19734