Saved in:
| Main Authors: | Lubana, Ekdeep Singh, Rager, Can, Hindupur, Sai Sumedh R., Costa, Valerie, Tuckute, Greta, Patel, Oam, Murthy, Sonia Krishna, Fel, Thomas, Wurgaft, Daniel, Bigelow, Eric J., Lin, Johnny, Ba, Demba, Wattenberg, Martin, Viegas, Fernanda, Weber, Melanie, Mueller, Aaron |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.01836 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
by: Hindupur, Sai Sumedh R., et al.
Published: (2025)
by: Hindupur, Sai Sumedh R., et al.
Published: (2025)
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
by: Costa, Valérie, et al.
Published: (2025)
by: Costa, Valérie, et al.
Published: (2025)
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
by: Costa, Valérie, et al.
Published: (2025)
by: Costa, Valérie, et al.
Published: (2025)
Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
by: Grosso, Gaia, et al.
Published: (2025)
by: Grosso, Gaia, et al.
Published: (2025)
Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models
by: Fel, Thomas, et al.
Published: (2025)
by: Fel, Thomas, et al.
Published: (2025)
Clustering Inductive Biases with Unrolled Networks
by: Huml, Jonathan, et al.
Published: (2023)
by: Huml, Jonathan, et al.
Published: (2023)
Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
by: Fel, Thomas, et al.
Published: (2025)
by: Fel, Thomas, et al.
Published: (2025)
Do Sparse Autoencoders Capture Concept Manifolds?
by: Bhalla, Usha, et al.
Published: (2026)
by: Bhalla, Usha, et al.
Published: (2026)
Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
by: Feucht, Sheridan, et al.
Published: (2026)
by: Feucht, Sheridan, et al.
Published: (2026)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
by: Bigelow, Eric, et al.
Published: (2025)
by: Bigelow, Eric, et al.
Published: (2025)
Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
by: Wurgaft, Daniel, et al.
Published: (2026)
by: Wurgaft, Daniel, et al.
Published: (2026)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
by: Zur, Amir, et al.
Published: (2025)
by: Zur, Amir, et al.
Published: (2025)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
by: Bohacek, Matyas, et al.
Published: (2025)
by: Bohacek, Matyas, et al.
Published: (2025)
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
by: Bigelow, Eric, et al.
Published: (2026)
by: Bigelow, Eric, et al.
Published: (2026)
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
by: Li, Kenneth, et al.
Published: (2023)
by: Li, Kenneth, et al.
Published: (2023)
The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
by: Sarfati, Raphaël, et al.
Published: (2026)
by: Sarfati, Raphaël, et al.
Published: (2026)
In-Context Learning Strategies Emerge Rationally
by: Wurgaft, Daniel, et al.
Published: (2025)
by: Wurgaft, Daniel, et al.
Published: (2025)
In-Context Learning Dynamics with Random Binary Sequences
by: Bigelow, Eric J., et al.
Published: (2023)
by: Bigelow, Eric J., et al.
Published: (2023)
Model Connectomes: A Generational Approach to Data-Efficient Language Models
by: Kotar, Klemen, et al.
Published: (2025)
by: Kotar, Klemen, et al.
Published: (2025)
Block-Recurrent Dynamics in Vision Transformers
by: Jacobs, Mozes, et al.
Published: (2025)
by: Jacobs, Mozes, et al.
Published: (2025)
Emergence of Hierarchical Emotion Organization in Large Language Models
by: Zhao, Bo, et al.
Published: (2025)
by: Zhao, Bo, et al.
Published: (2025)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
by: Jaipersaud, Brandon, et al.
Published: (2025)
by: Jaipersaud, Brandon, et al.
Published: (2025)
Abrupt Learning in Transformers: A Case Study on Matrix Completion
by: Gopalani, Pulkit, et al.
Published: (2024)
by: Gopalani, Pulkit, et al.
Published: (2024)
Relational Composition in Neural Networks: A Survey and Call to Action
by: Wattenberg, Martin, et al.
Published: (2024)
by: Wattenberg, Martin, et al.
Published: (2024)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Topoformer: brain-like topographic organization in Transformer language models through spatial querying and reweighting
by: Binhuraib, Taha, et al.
Published: (2025)
by: Binhuraib, Taha, et al.
Published: (2025)
ICLR: In-Context Learning of Representations
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
What Does it Mean for a Neural Network to Learn a "World Model"?
by: Li, Kenneth, et al.
Published: (2025)
by: Li, Kenneth, et al.
Published: (2025)
Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions
by: Lee, Andrew, et al.
Published: (2026)
by: Lee, Andrew, et al.
Published: (2026)
Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
by: Huang, Jing, et al.
Published: (2026)
by: Huang, Jing, et al.
Published: (2026)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
by: Pres, Itamar, et al.
Published: (2024)
by: Pres, Itamar, et al.
Published: (2024)
Analyzing (In)Abilities of SAEs via Formal Languages
by: Menon, Abhinav, et al.
Published: (2024)
by: Menon, Abhinav, et al.
Published: (2024)
Implicit Generative Modeling by Kernel Similarity Matching
by: Choudhary, Shubham, et al.
Published: (2025)
by: Choudhary, Shubham, et al.
Published: (2025)
Weighed l1 on the simplex: Compressive sensing meets locality
by: Tasissa, Abiy, et al.
Published: (2021)
by: Tasissa, Abiy, et al.
Published: (2021)
A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language
by: Lubana, Ekdeep Singh, et al.
Published: (2024)
by: Lubana, Ekdeep Singh, et al.
Published: (2024)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
by: Okawa, Maya, et al.
Published: (2023)
by: Okawa, Maya, et al.
Published: (2023)
Dialogue Action Tokens: Steering Language Models in Goal-Directed Dialogue with a Multi-Turn Planner
by: Li, Kenneth, et al.
Published: (2024)
by: Li, Kenneth, et al.
Published: (2024)
Decomposing Query-Key Feature Interactions Using Contrastive Covariances
by: Lee, Andrew, et al.
Published: (2026)
by: Lee, Andrew, et al.
Published: (2026)
When Bad Data Leads to Good Models
by: Li, Kenneth, et al.
Published: (2025)
by: Li, Kenneth, et al.
Published: (2025)
Similar Items
-
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
by: Hindupur, Sai Sumedh R., et al.
Published: (2025) -
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
by: Costa, Valérie, et al.
Published: (2025) -
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
by: Costa, Valérie, et al.
Published: (2025) -
Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
by: Grosso, Gaia, et al.
Published: (2025) -
Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models
by: Fel, Thomas, et al.
Published: (2025)