Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hindupur, Sai Sumedh R., Lubana, Ekdeep Singh, Fel, Thomas, Ba, Demba |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
von: Grosso, Gaia, et al.
Veröffentlicht: (2025)
von: Grosso, Gaia, et al.
Veröffentlicht: (2025)
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
von: Costa, Valérie, et al.
Veröffentlicht: (2025)
von: Costa, Valérie, et al.
Veröffentlicht: (2025)
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
von: Costa, Valérie, et al.
Veröffentlicht: (2025)
von: Costa, Valérie, et al.
Veröffentlicht: (2025)
Do Sparse Autoencoders Capture Concept Manifolds?
von: Bhalla, Usha, et al.
Veröffentlicht: (2026)
von: Bhalla, Usha, et al.
Veröffentlicht: (2026)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
von: Bohacek, Matyas, et al.
Veröffentlicht: (2025)
von: Bohacek, Matyas, et al.
Veröffentlicht: (2025)
Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
von: Fel, Thomas, et al.
Veröffentlicht: (2025)
von: Fel, Thomas, et al.
Veröffentlicht: (2025)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2025)
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2025)
A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language
von: Lubana, Ekdeep Singh, et al.
Veröffentlicht: (2024)
von: Lubana, Ekdeep Singh, et al.
Veröffentlicht: (2024)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
Block-Recurrent Dynamics in Vision Transformers
von: Jacobs, Mozes, et al.
Veröffentlicht: (2025)
von: Jacobs, Mozes, et al.
Veröffentlicht: (2025)
Priors in Time: Missing Inductive Biases for Language Model Interpretability
von: Lubana, Ekdeep Singh, et al.
Veröffentlicht: (2025)
von: Lubana, Ekdeep Singh, et al.
Veröffentlicht: (2025)
The Geometry of Concepts: Sparse Autoencoder Feature Structure
von: Li, Yuxiao, et al.
Veröffentlicht: (2024)
von: Li, Yuxiao, et al.
Veröffentlicht: (2024)
The Impact of Off-Policy Training Data on Probe Generalisation
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
Clustering Inductive Biases with Unrolled Networks
von: Huml, Jonathan, et al.
Veröffentlicht: (2023)
von: Huml, Jonathan, et al.
Veröffentlicht: (2023)
In-Context Learning Strategies Emerge Rationally
von: Wurgaft, Daniel, et al.
Veröffentlicht: (2025)
von: Wurgaft, Daniel, et al.
Veröffentlicht: (2025)
In-Context Learning Dynamics with Random Binary Sequences
von: Bigelow, Eric J., et al.
Veröffentlicht: (2023)
von: Bigelow, Eric J., et al.
Veröffentlicht: (2023)
Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation Model
von: Khona, Mikail, et al.
Veröffentlicht: (2024)
von: Khona, Mikail, et al.
Veröffentlicht: (2024)
A Geometric Unification of Concept Learning with Concept Cones
von: Rocchi--Henry, Alexandre, et al.
Veröffentlicht: (2025)
von: Rocchi--Henry, Alexandre, et al.
Veröffentlicht: (2025)
Finding Belief Geometries with Sparse Autoencoders
von: Levinson, Matthew
Veröffentlicht: (2026)
von: Levinson, Matthew
Veröffentlicht: (2026)
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
von: Bigelow, Eric, et al.
Veröffentlicht: (2026)
von: Bigelow, Eric, et al.
Veröffentlicht: (2026)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
von: Bigelow, Eric, et al.
Veröffentlicht: (2025)
von: Bigelow, Eric, et al.
Veröffentlicht: (2025)
Emergence of Hierarchical Emotion Organization in Large Language Models
von: Zhao, Bo, et al.
Veröffentlicht: (2025)
von: Zhao, Bo, et al.
Veröffentlicht: (2025)
SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
ICLR: In-Context Learning of Representations
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
Sparse Autoencoders, Again?
von: Lu, Yin, et al.
Veröffentlicht: (2025)
von: Lu, Yin, et al.
Veröffentlicht: (2025)
Rethinking Sparse Autoencoders: Select-and-Project for Fairness and Control from Encoder Features Alone
von: Bărbălau, Antonio, et al.
Veröffentlicht: (2025)
von: Bărbălau, Antonio, et al.
Veröffentlicht: (2025)
Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts
von: Peng, Kenny, et al.
Veröffentlicht: (2025)
von: Peng, Kenny, et al.
Veröffentlicht: (2025)
Policy-Guided Search on Tree-of-Thoughts for Efficient Problem Solving with Bounded Language Model Queries
von: Pendurkar, Sumedh, et al.
Veröffentlicht: (2026)
von: Pendurkar, Sumedh, et al.
Veröffentlicht: (2026)
Are Sparse Autoencoder Benchmarks Reliable?
von: Chanin, David
Veröffentlicht: (2026)
von: Chanin, David
Veröffentlicht: (2026)
Abrupt Learning in Transformers: A Case Study on Matrix Completion
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2024)
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2024)
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
von: Joshi, Shruti, et al.
Veröffentlicht: (2025)
von: Joshi, Shruti, et al.
Veröffentlicht: (2025)
BatchTopK Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
Improving Sparse Autoencoder with Dynamic Attention
von: Wang, Dongsheng, et al.
Veröffentlicht: (2026)
von: Wang, Dongsheng, et al.
Veröffentlicht: (2026)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2024)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2024)
One Wave To Explain Them All: A Unifying Perspective On Feature Attribution
von: Kasmi, Gabriel, et al.
Veröffentlicht: (2024)
von: Kasmi, Gabriel, et al.
Veröffentlicht: (2024)
Sparse Autoencoder Features for Classifications and Transferability
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
von: Ayonrinde, Kola
Veröffentlicht: (2024)
von: Ayonrinde, Kola
Veröffentlicht: (2024)
Ähnliche Einträge
-
Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
von: Grosso, Gaia, et al.
Veröffentlicht: (2025) -
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
von: Costa, Valérie, et al.
Veröffentlicht: (2025) -
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
von: Costa, Valérie, et al.
Veröffentlicht: (2025) -
Do Sparse Autoencoders Capture Concept Manifolds?
von: Bhalla, Usha, et al.
Veröffentlicht: (2026) -
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
von: Bohacek, Matyas, et al.
Veröffentlicht: (2025)