Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Fel, Thomas, Lubana, Ekdeep Singh, Prince, Jacob S., Kowal, Matthew, Boutin, Victor, Papadimitriou, Isabel, Wang, Binxu, Wattenberg, Martin, Ba, Demba, Konkle, Talia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
di: Fel, Thomas, et al.
Pubblicazione: (2025)
di: Fel, Thomas, et al.
Pubblicazione: (2025)
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
di: Hindupur, Sai Sumedh R., et al.
Pubblicazione: (2025)
di: Hindupur, Sai Sumedh R., et al.
Pubblicazione: (2025)
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
di: Costa, Valérie, et al.
Pubblicazione: (2025)
di: Costa, Valérie, et al.
Pubblicazione: (2025)
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
di: Costa, Valérie, et al.
Pubblicazione: (2025)
di: Costa, Valérie, et al.
Pubblicazione: (2025)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
di: Bohacek, Matyas, et al.
Pubblicazione: (2025)
di: Bohacek, Matyas, et al.
Pubblicazione: (2025)
Bi-Orthogonal Factor Decomposition for Vision Transformers
di: Doshi, Fenil R., et al.
Pubblicazione: (2026)
di: Doshi, Fenil R., et al.
Pubblicazione: (2026)
Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models
di: Doshi, Fenil R., et al.
Pubblicazione: (2025)
di: Doshi, Fenil R., et al.
Pubblicazione: (2025)
Do Sparse Autoencoders Capture Concept Manifolds?
di: Bhalla, Usha, et al.
Pubblicazione: (2026)
di: Bhalla, Usha, et al.
Pubblicazione: (2026)
Vocabulary embeddings organize linguistic structure early in language model training
di: Papadimitriou, Isabel, et al.
Pubblicazione: (2025)
di: Papadimitriou, Isabel, et al.
Pubblicazione: (2025)
Block-Recurrent Dynamics in Vision Transformers
di: Jacobs, Mozes, et al.
Pubblicazione: (2025)
di: Jacobs, Mozes, et al.
Pubblicazione: (2025)
Feature Accentuation: Revealing 'What' Features Respond to in Natural Images
di: Hamblin, Chris, et al.
Pubblicazione: (2024)
di: Hamblin, Chris, et al.
Pubblicazione: (2024)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2025)
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2025)
Abrupt Learning in Transformers: A Case Study on Matrix Completion
di: Gopalani, Pulkit, et al.
Pubblicazione: (2024)
di: Gopalani, Pulkit, et al.
Pubblicazione: (2024)
K-Deep Simplex: Deep Manifold Learning via Local Dictionaries
di: Tankala, Pranay, et al.
Pubblicazione: (2020)
di: Tankala, Pranay, et al.
Pubblicazione: (2020)
Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
Priors in Time: Missing Inductive Biases for Language Model Interpretability
di: Lubana, Ekdeep Singh, et al.
Pubblicazione: (2025)
di: Lubana, Ekdeep Singh, et al.
Pubblicazione: (2025)
Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
di: Thasarathan, Harrish, et al.
Pubblicazione: (2025)
di: Thasarathan, Harrish, et al.
Pubblicazione: (2025)
Swing-by Dynamics in Concept Learning and Compositional Generalization
di: Yang, Yongyi, et al.
Pubblicazione: (2024)
di: Yang, Yongyi, et al.
Pubblicazione: (2024)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
di: Zur, Amir, et al.
Pubblicazione: (2025)
di: Zur, Amir, et al.
Pubblicazione: (2025)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
di: Pres, Itamar, et al.
Pubblicazione: (2024)
di: Pres, Itamar, et al.
Pubblicazione: (2024)
Analyzing (In)Abilities of SAEs via Formal Languages
di: Menon, Abhinav, et al.
Pubblicazione: (2024)
di: Menon, Abhinav, et al.
Pubblicazione: (2024)
ICLR: In-Context Learning of Representations
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings
di: Dhimoïla, Grégoire, et al.
Pubblicazione: (2026)
di: Dhimoïla, Grégoire, et al.
Pubblicazione: (2026)
Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
di: Feucht, Sheridan, et al.
Pubblicazione: (2026)
di: Feucht, Sheridan, et al.
Pubblicazione: (2026)
A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language
di: Lubana, Ekdeep Singh, et al.
Pubblicazione: (2024)
di: Lubana, Ekdeep Singh, et al.
Pubblicazione: (2024)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
di: Okawa, Maya, et al.
Pubblicazione: (2023)
di: Okawa, Maya, et al.
Pubblicazione: (2023)
Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
di: Grosso, Gaia, et al.
Pubblicazione: (2025)
di: Grosso, Gaia, et al.
Pubblicazione: (2025)
Clustering Inductive Biases with Unrolled Networks
di: Huml, Jonathan, et al.
Pubblicazione: (2023)
di: Huml, Jonathan, et al.
Pubblicazione: (2023)
Implicit Generative Modeling by Kernel Similarity Matching
di: Choudhary, Shubham, et al.
Pubblicazione: (2025)
di: Choudhary, Shubham, et al.
Pubblicazione: (2025)
Weighed l1 on the simplex: Compressive sensing meets locality
di: Tasissa, Abiy, et al.
Pubblicazione: (2021)
di: Tasissa, Abiy, et al.
Pubblicazione: (2021)
Interpreting the linear structure of vision-language model embedding spaces
di: Papadimitriou, Isabel, et al.
Pubblicazione: (2025)
di: Papadimitriou, Isabel, et al.
Pubblicazione: (2025)
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks
di: Ramesh, Rahul, et al.
Pubblicazione: (2023)
di: Ramesh, Rahul, et al.
Pubblicazione: (2023)
Understanding Inhibition Through Maximally Tense Images
di: Hamblin, Chris, et al.
Pubblicazione: (2024)
di: Hamblin, Chris, et al.
Pubblicazione: (2024)
FOVI: A biologically-inspired foveated interface for deep vision models
di: Blauch, Nicholas M., et al.
Pubblicazione: (2026)
di: Blauch, Nicholas M., et al.
Pubblicazione: (2026)
Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
di: Wurgaft, Daniel, et al.
Pubblicazione: (2026)
di: Wurgaft, Daniel, et al.
Pubblicazione: (2026)
Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing
di: Nishi, Kento, et al.
Pubblicazione: (2024)
di: Nishi, Kento, et al.
Pubblicazione: (2024)
The Impact of Off-Policy Training Data on Probe Generalisation
di: Kirch, Nathalie, et al.
Pubblicazione: (2025)
di: Kirch, Nathalie, et al.
Pubblicazione: (2025)
Sparks of Explainability: Recent Advancements in Explaining Large Vision Models
di: Fel, Thomas
Pubblicazione: (2025)
di: Fel, Thomas
Pubblicazione: (2025)
Documenti analoghi
-
Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
di: Fel, Thomas, et al.
Pubblicazione: (2025) -
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
di: Hindupur, Sai Sumedh R., et al.
Pubblicazione: (2025) -
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
di: Costa, Valérie, et al.
Pubblicazione: (2025) -
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
di: Costa, Valérie, et al.
Pubblicazione: (2025) -
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
di: Bohacek, Matyas, et al.
Pubblicazione: (2025)