Scaling and evaluating sparse autoencoders
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Leo, la Tour, Tom Dupré, Tillman, Henk, Goh, Gabriel, Troll, Rajan, Radford, Alec, Sutskever, Ilya, Leike, Jan, Wu, Jeffrey |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Investigating task-specific prompts and sparse autoencoders for activation monitoring
by: Tillman, Henk, et al.
Published: (2025)
by: Tillman, Henk, et al.
Published: (2025)
Faster and Simpler Greedy Algorithm for $k$-Median and $k$-Means
by: la Tour, Max Dupré, et al.
Published: (2024)
by: la Tour, Max Dupré, et al.
Published: (2024)
Bad News for Couples: Tight Lower Bounds for Fair Division of Indivisible Items
by: la Tour, Max Dupré
Published: (2026)
by: la Tour, Max Dupré
Published: (2026)
Discrepancy And Fair Division For Non-Additive Valuations
by: la Tour, Max Dupre, et al.
Published: (2025)
by: la Tour, Max Dupre, et al.
Published: (2025)
Gerrymandering Planar Graphs
by: Dippel, Jack, et al.
Published: (2023)
by: Dippel, Jack, et al.
Published: (2023)
Shaping capabilities with token-level data filtering
by: Rathi, Neil, et al.
Published: (2026)
by: Rathi, Neil, et al.
Published: (2026)
Applying sparse autoencoders to unlearn knowledge in language models
by: Farrell, Eoin, et al.
Published: (2024)
by: Farrell, Eoin, et al.
Published: (2024)
Steering CLIP's vision transformer with sparse autoencoders
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
Recognizing Leaf Powers and Pairwise Compatibility Graphs is NP-Complete
by: la Tour, Max Dupré, et al.
Published: (2025)
by: la Tour, Max Dupré, et al.
Published: (2025)
On the hardness of recognizing graphs of small mim-width and its variants
by: la Tour, Max Dupré, et al.
Published: (2025)
by: la Tour, Max Dupré, et al.
Published: (2025)
Making Old Things New: A Unified Algorithm for Differentially Private Clustering
by: la Tour, Max Dupré, et al.
Published: (2024)
by: la Tour, Max Dupré, et al.
Published: (2024)
Fully Dynamic k-Means Coreset in Near-Optimal Update Time
by: la Tour, Max Dupré, et al.
Published: (2024)
by: la Tour, Max Dupré, et al.
Published: (2024)
Can sparse autoencoders be used to decompose and interpret steering vectors?
by: Mayne, Harry, et al.
Published: (2024)
by: Mayne, Harry, et al.
Published: (2024)
Persona Features Control Emergent Misalignment
by: Wang, Miles, et al.
Published: (2025)
by: Wang, Miles, et al.
Published: (2025)
Excess Description Length of Learning Generalizable Predictors
by: Donoway, Elizabeth, et al.
Published: (2026)
by: Donoway, Elizabeth, et al.
Published: (2026)
Transformer autoencoder with local attention for sparse and irregular time series with application on risk estimation
by: Rodis, Panteleimon
Published: (2026)
by: Rodis, Panteleimon
Published: (2026)
$k$-Leaf Powers Cannot be Characterized by a Finite Set of Forbidden Induced Subgraphs for $k \geq 5$
by: la Tour, Max Dupré, et al.
Published: (2024)
by: la Tour, Max Dupré, et al.
Published: (2024)
Understanding sparse autoencoder scaling in the presence of feature manifolds
by: Michaud, Eric J., et al.
Published: (2025)
by: Michaud, Eric J., et al.
Published: (2025)
Tight Asymptotic Bounds for Fair Division With Externalities
by: Connor, Frank, et al.
Published: (2026)
by: Connor, Frank, et al.
Published: (2026)
Weight-sparse transformers have interpretable circuits
by: Gao, Leo, et al.
Published: (2025)
by: Gao, Leo, et al.
Published: (2025)
Learning a Generative Meta-Model of LLM Activations
by: Luo, Grace, et al.
Published: (2026)
by: Luo, Grace, et al.
Published: (2026)
Eliminating Majority Illusion is Easy
by: Dippel, Jack, et al.
Published: (2024)
by: Dippel, Jack, et al.
Published: (2024)
Ecología del paisaje
by: Carl Troll
Published: (2003)
by: Carl Troll
Published: (2003)
GRU‐based stacked sparse autoencoder with attention mechanism for process monitoring
by: Zengdi Miao, et al.
Published: (2025)
by: Zengdi Miao, et al.
Published: (2025)
Scaling sparse feature circuit finding for in-context learning
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
Proportion and Perspective Control for Flow-Based Image Generation
by: Boudier, Julien, et al.
Published: (2025)
by: Boudier, Julien, et al.
Published: (2025)
Quality‐related process monitoring approach based on sparse autoencoder and comprehensive KPLS
by: Yikai Xue, et al.
Published: (2025)
by: Yikai Xue, et al.
Published: (2025)
Decomposing multimodal embedding spaces with group-sparse autoencoders
by: Kaushik, Chiraag, et al.
Published: (2026)
by: Kaushik, Chiraag, et al.
Published: (2026)
Analysis of chaotic dynamical systems with autoencoders
by: Almazova, N., et al.
Published: (2021)
by: Almazova, N., et al.
Published: (2021)
Variational autoencoder-based neural network model compression
by: Cheng, Liang, et al.
Published: (2024)
by: Cheng, Liang, et al.
Published: (2024)
Multi-view autoencoders for Fake News Detection
by: Pereira, Ingryd V. S. T., et al.
Published: (2025)
by: Pereira, Ingryd V. S. T., et al.
Published: (2025)
Discretization of continuous input spaces in the hippocampal autoencoder
by: Amil, Adrian F., et al.
Published: (2024)
by: Amil, Adrian F., et al.
Published: (2024)
Usage and Usability Assessment: Library Practices and Concerns.
by: Covey, Denise Troll
Published: (2002)
by: Covey, Denise Troll
Published: (2002)
Information Technologies at Carnegie Mellon.
by: Troll, Denise A.
Published: (1992)
by: Troll, Denise A.
Published: (1992)
How and Why Libraries Are Changing: What We Know and What We Need To Know.
by: Troll, Denise A.
Published: (2002)
by: Troll, Denise A.
Published: (2002)
Exploring Content and Social Connections of Fake News with Explainable Text and Graph Learning
by: Lourenço, Vítor N., et al.
Published: (2025)
by: Lourenço, Vítor N., et al.
Published: (2025)
Multimodal normative modeling in Alzheimers Disease with introspective variational autoencoders
by: Kumar, Sayantan, et al.
Published: (2026)
by: Kumar, Sayantan, et al.
Published: (2026)
OWL2Vec4OA: Tailoring Knowledge Graph Embeddings for Ontology Alignment
by: Teymurova, Sevinj, et al.
Published: (2024)
by: Teymurova, Sevinj, et al.
Published: (2024)
Metabolic cost of information processing in Poisson variational autoencoders
by: Vafaii, Hadi, et al.
Published: (2026)
by: Vafaii, Hadi, et al.
Published: (2026)
Hybrid guided variational autoencoder for visual place recognition
by: Wang, Ni, et al.
Published: (2026)
by: Wang, Ni, et al.
Published: (2026)
Similar Items
-
Investigating task-specific prompts and sparse autoencoders for activation monitoring
by: Tillman, Henk, et al.
Published: (2025) -
Faster and Simpler Greedy Algorithm for $k$-Median and $k$-Means
by: la Tour, Max Dupré, et al.
Published: (2024) -
Bad News for Couples: Tight Lower Bounds for Fair Division of Indivisible Items
by: la Tour, Max Dupré
Published: (2026) -
Discrepancy And Fair Division For Non-Additive Valuations
by: la Tour, Max Dupre, et al.
Published: (2025) -
Gerrymandering Planar Graphs
by: Dippel, Jack, et al.
Published: (2023)