Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Kenny, Movva, Rajiv, Kleinberg, Jon, Pierson, Emma, Garg, Nikhil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sparse Autoencoders for Hypothesis Generation
by: Movva, Rajiv, et al.
Published: (2025)
by: Movva, Rajiv, et al.
Published: (2025)
How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
by: Garg, Nikhil, et al.
Published: (2026)
by: Garg, Nikhil, et al.
Published: (2026)
Correlated Errors in Large Language Models
by: Kim, Elliot, et al.
Published: (2025)
by: Kim, Elliot, et al.
Published: (2025)
Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers
by: Movva, Rajiv, et al.
Published: (2023)
by: Movva, Rajiv, et al.
Published: (2023)
A No Free Lunch Theorem for Human-AI Collaboration
by: Peng, Kenny, et al.
Published: (2024)
by: Peng, Kenny, et al.
Published: (2024)
Generative AI in Medicine
by: Shanmugam, Divya, et al.
Published: (2024)
by: Shanmugam, Divya, et al.
Published: (2024)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
by: Li, Aaron J., et al.
Published: (2025)
by: Li, Aaron J., et al.
Published: (2025)
What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
by: Movva, Rajiv, et al.
Published: (2025)
by: Movva, Rajiv, et al.
Published: (2025)
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
by: Joshi, Shruti, et al.
Published: (2025)
by: Joshi, Shruti, et al.
Published: (2025)
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
by: Muhamed, Aashiq, et al.
Published: (2024)
by: Muhamed, Aashiq, et al.
Published: (2024)
Learning Disease Progression Models That Capture Health Disparities
by: Chiang, Erica, et al.
Published: (2024)
by: Chiang, Erica, et al.
Published: (2024)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs
by: Li, Yinghui, et al.
Published: (2025)
by: Li, Yinghui, et al.
Published: (2025)
A Bayesian Spatial Model to Correct Under-Reporting in Urban Crowdsourcing
by: Agostini, Gabriel, et al.
Published: (2023)
by: Agostini, Gabriel, et al.
Published: (2023)
Position: The Most Expensive Part of an LLM should be its Training Data
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
by: Huang, Saffron, et al.
Published: (2025)
by: Huang, Saffron, et al.
Published: (2025)
Cross-Layer Discrete Concept Discovery for Interpreting Language Models
by: Garg, Ankur, et al.
Published: (2025)
by: Garg, Ankur, et al.
Published: (2025)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
by: Guldimann, Philipp, et al.
Published: (2024)
by: Guldimann, Philipp, et al.
Published: (2024)
Emissions and Performance Trade-off Between Small and Large Language Models
by: Garg, Anandita, et al.
Published: (2025)
by: Garg, Anandita, et al.
Published: (2025)
Language Generation in the Limit
by: Kleinberg, Jon, et al.
Published: (2024)
by: Kleinberg, Jon, et al.
Published: (2024)
An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt Learning
by: Jin, Chen, et al.
Published: (2023)
by: Jin, Chen, et al.
Published: (2023)
Anatomy of a Machine Learning Ecosystem: 2 Million Models on Hugging Face
by: Laufer, Benjamin, et al.
Published: (2025)
by: Laufer, Benjamin, et al.
Published: (2025)
Designing Skill-Compatible AI: Methodologies and Frameworks in Chess
by: Hamade, Karim, et al.
Published: (2024)
by: Hamade, Karim, et al.
Published: (2024)
Concept Tokens: Learning Behavioral Embeddings Through Concept Definitions
by: Sastre, Ignacio, et al.
Published: (2026)
by: Sastre, Ignacio, et al.
Published: (2026)
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
by: Li, Nathaniel, et al.
Published: (2024)
by: Li, Nathaniel, et al.
Published: (2024)
Bayesian Modeling of Zero-Shot Classifications for Urban Flood Detection
by: Franchi, Matt, et al.
Published: (2025)
by: Franchi, Matt, et al.
Published: (2025)
Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
by: Saha, Rounak, et al.
Published: (2026)
by: Saha, Rounak, et al.
Published: (2026)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
by: Su, Jingtong, et al.
Published: (2025)
by: Su, Jingtong, et al.
Published: (2025)
Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution
by: Zhao, Haiyan, et al.
Published: (2024)
by: Zhao, Haiyan, et al.
Published: (2024)
Sparse Autoencoder Features for Classifications and Transferability
by: Gallifant, Jack, et al.
Published: (2025)
by: Gallifant, Jack, et al.
Published: (2025)
LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases
by: Bouchard, Dylan, et al.
Published: (2025)
by: Bouchard, Dylan, et al.
Published: (2025)
Psychological Counseling Ability of Large Language Models
by: Peng, Fangyu, et al.
Published: (2025)
by: Peng, Fangyu, et al.
Published: (2025)
ICE-T: A Multi-Faceted Concept for Teaching Machine Learning
by: Krone, Hendrik, et al.
Published: (2024)
by: Krone, Hendrik, et al.
Published: (2024)
Safe and Certifiable AI Systems: Concepts, Challenges, and Lessons Learned
by: Schweighofer, Kajetan, et al.
Published: (2025)
by: Schweighofer, Kajetan, et al.
Published: (2025)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
by: Muchane, Mark, et al.
Published: (2025)
by: Muchane, Mark, et al.
Published: (2025)
Exploring Concept Depth: How Large Language Models Acquire Knowledge and Concept at Different Layers?
by: Jin, Mingyu, et al.
Published: (2024)
by: Jin, Mingyu, et al.
Published: (2024)
AlignSAE: Concept-Aligned Sparse Autoencoders
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
PICKT: Practical Interlinked Concept Knowledge Tracing for Personalized Learning using Knowledge Map Concept Relations
by: Lee, Wonbeen, et al.
Published: (2025)
by: Lee, Wonbeen, et al.
Published: (2025)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
by: Chanin, David, et al.
Published: (2025)
by: Chanin, David, et al.
Published: (2025)
Disentangling Heterogeneous Knowledge Concept Embedding for Cognitive Diagnosis on Untested Knowledge
by: Zhang, Miao, et al.
Published: (2024)
by: Zhang, Miao, et al.
Published: (2024)
Similar Items
-
Sparse Autoencoders for Hypothesis Generation
by: Movva, Rajiv, et al.
Published: (2025) -
How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
by: Garg, Nikhil, et al.
Published: (2026) -
Correlated Errors in Large Language Models
by: Kim, Elliot, et al.
Published: (2025) -
Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers
by: Movva, Rajiv, et al.
Published: (2023) -
A No Free Lunch Theorem for Human-AI Collaboration
by: Peng, Kenny, et al.
Published: (2024)