Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
Fuente:
arXiv
Saved in:
| Main Authors: | Fel, Thomas, Wang, Binxu, Lepori, Michael A., Kowal, Matthew, Lee, Andrew, Balestriero, Randall, Joseph, Sonia, Lubana, Ekdeep S., Konkle, Talia, Ba, Demba, Wattenberg, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models
by: Fel, Thomas, et al.
Published: (2025)
by: Fel, Thomas, et al.
Published: (2025)
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
by: Hindupur, Sai Sumedh R., et al.
Published: (2025)
by: Hindupur, Sai Sumedh R., et al.
Published: (2025)
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
by: Costa, Valérie, et al.
Published: (2025)
by: Costa, Valérie, et al.
Published: (2025)
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
by: Costa, Valérie, et al.
Published: (2025)
by: Costa, Valérie, et al.
Published: (2025)
Interpreting Physics in Video World Models
by: Joseph, Sonia, et al.
Published: (2026)
by: Joseph, Sonia, et al.
Published: (2026)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
by: Bohacek, Matyas, et al.
Published: (2025)
by: Bohacek, Matyas, et al.
Published: (2025)
Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models
by: Doshi, Fenil R., et al.
Published: (2025)
by: Doshi, Fenil R., et al.
Published: (2025)
Bi-Orthogonal Factor Decomposition for Vision Transformers
by: Doshi, Fenil R., et al.
Published: (2026)
by: Doshi, Fenil R., et al.
Published: (2026)
Do Sparse Autoencoders Capture Concept Manifolds?
by: Bhalla, Usha, et al.
Published: (2026)
by: Bhalla, Usha, et al.
Published: (2026)
Feature Accentuation: Revealing 'What' Features Respond to in Natural Images
by: Hamblin, Chris, et al.
Published: (2024)
by: Hamblin, Chris, et al.
Published: (2024)
Priors in Time: Missing Inductive Biases for Language Model Interpretability
by: Lubana, Ekdeep Singh, et al.
Published: (2025)
by: Lubana, Ekdeep Singh, et al.
Published: (2025)
Task Priors: Enhancing Model Evaluation by Considering the Entire Space of Downstream Tasks
by: Patel, Niket, et al.
Published: (2025)
by: Patel, Niket, et al.
Published: (2025)
Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
by: Thasarathan, Harrish, et al.
Published: (2025)
by: Thasarathan, Harrish, et al.
Published: (2025)
Block-Recurrent Dynamics in Vision Transformers
by: Jacobs, Mozes, et al.
Published: (2025)
by: Jacobs, Mozes, et al.
Published: (2025)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
by: Okawa, Maya, et al.
Published: (2023)
by: Okawa, Maya, et al.
Published: (2023)
Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
by: Wurgaft, Daniel, et al.
Published: (2026)
by: Wurgaft, Daniel, et al.
Published: (2026)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
by: Jaipersaud, Brandon, et al.
Published: (2025)
by: Jaipersaud, Brandon, et al.
Published: (2025)
Abrupt Learning in Transformers: A Case Study on Matrix Completion
by: Gopalani, Pulkit, et al.
Published: (2024)
by: Gopalani, Pulkit, et al.
Published: (2024)
ICLR: In-Context Learning of Representations
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
On the Geometry of Deep Learning
by: Balestriero, Randall, et al.
Published: (2024)
by: Balestriero, Randall, et al.
Published: (2024)
Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation
by: Balestriero, Randall, et al.
Published: (2023)
by: Balestriero, Randall, et al.
Published: (2023)
Swing-by Dynamics in Concept Learning and Compositional Generalization
by: Yang, Yongyi, et al.
Published: (2024)
by: Yang, Yongyi, et al.
Published: (2024)
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks
by: Ramesh, Rahul, et al.
Published: (2023)
by: Ramesh, Rahul, et al.
Published: (2023)
Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
by: Grosso, Gaia, et al.
Published: (2025)
by: Grosso, Gaia, et al.
Published: (2025)
The Geometry of Self-Verification in a Task-Specific Reasoning Model
by: Lee, Andrew, et al.
Published: (2025)
by: Lee, Andrew, et al.
Published: (2025)
Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
by: Feucht, Sheridan, et al.
Published: (2026)
by: Feucht, Sheridan, et al.
Published: (2026)
Clustering Inductive Biases with Unrolled Networks
by: Huml, Jonathan, et al.
Published: (2023)
by: Huml, Jonathan, et al.
Published: (2023)
Implicit Generative Modeling by Kernel Similarity Matching
by: Choudhary, Shubham, et al.
Published: (2025)
by: Choudhary, Shubham, et al.
Published: (2025)
Weighed l1 on the simplex: Compressive sensing meets locality
by: Tasissa, Abiy, et al.
Published: (2021)
by: Tasissa, Abiy, et al.
Published: (2021)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
by: Zur, Amir, et al.
Published: (2025)
by: Zur, Amir, et al.
Published: (2025)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
by: Pres, Itamar, et al.
Published: (2024)
by: Pres, Itamar, et al.
Published: (2024)
Analyzing (In)Abilities of SAEs via Formal Languages
by: Menon, Abhinav, et al.
Published: (2024)
by: Menon, Abhinav, et al.
Published: (2024)
No Location Left Behind: Measuring and Improving the Fairness of Implicit Representations for Earth Data
by: Cai, Daniel, et al.
Published: (2025)
by: Cai, Daniel, et al.
Published: (2025)
Eidetic Learning: an Efficient and Provable Solution to Catastrophic Forgetting
by: Dronen, Nicholas, et al.
Published: (2025)
by: Dronen, Nicholas, et al.
Published: (2025)
SAFE: A Novel Approach to AI Weather Evaluation through Stratified Assessments of Forecasts over Earth
by: Masi, Nick, et al.
Published: (2025)
by: Masi, Nick, et al.
Published: (2025)
ALLoRA: Adaptive Learning Rate Mitigates LoRA Fatal Flaws
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
Understanding Inhibition Through Maximally Tense Images
by: Hamblin, Chris, et al.
Published: (2024)
by: Hamblin, Chris, et al.
Published: (2024)
FOVI: A biologically-inspired foveated interface for deep vision models
by: Blauch, Nicholas M., et al.
Published: (2026)
by: Blauch, Nicholas M., et al.
Published: (2026)
Similar Items
-
Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models
by: Fel, Thomas, et al.
Published: (2025) -
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
by: Hindupur, Sai Sumedh R., et al.
Published: (2025) -
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
by: Costa, Valérie, et al.
Published: (2025) -
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
by: Costa, Valérie, et al.
Published: (2025) -
Interpreting Physics in Video World Models
by: Joseph, Sonia, et al.
Published: (2026)