Efficient Dictionary Learning with Switch Sparse Autoencoders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mudide, Anish, Engels, Joshua, Michaud, Eric J., Tegmark, Max, de Witt, Christian Schroeder |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Decomposing The Dark Matter of Sparse Autoencoders
von: Engels, Joshua, et al.
Veröffentlicht: (2024)
von: Engels, Joshua, et al.
Veröffentlicht: (2024)
Low-Rank Adapting Models for Sparse Autoencoders
von: Chen, Matthew, et al.
Veröffentlicht: (2025)
von: Chen, Matthew, et al.
Veröffentlicht: (2025)
The Geometry of Concepts: Sparse Autoencoder Feature Structure
von: Li, Yuxiao, et al.
Veröffentlicht: (2024)
von: Li, Yuxiao, et al.
Veröffentlicht: (2024)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
Not All Language Model Features Are One-Dimensionally Linear
von: Engels, Joshua, et al.
Veröffentlicht: (2024)
von: Engels, Joshua, et al.
Veröffentlicht: (2024)
Opening the AI black box: program synthesis via mechanistic interpretability
von: Michaud, Eric J., et al.
Veröffentlicht: (2024)
von: Michaud, Eric J., et al.
Veröffentlicht: (2024)
On the creation of narrow AI: hierarchy and nonlocality of neural network skills
von: Michaud, Eric J., et al.
Veröffentlicht: (2025)
von: Michaud, Eric J., et al.
Veröffentlicht: (2025)
SAGE: Scalable Ground Truth Evaluations for Large Sparse Autoencoders
von: Venhoff, Constantin, et al.
Veröffentlicht: (2024)
von: Venhoff, Constantin, et al.
Veröffentlicht: (2024)
The Quantization Model of Neural Scaling
von: Michaud, Eric J., et al.
Veröffentlicht: (2023)
von: Michaud, Eric J., et al.
Veröffentlicht: (2023)
Physics of Skill Learning
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
Scaling Laws For Scalable Oversight
von: Engels, Joshua, et al.
Veröffentlicht: (2025)
von: Engels, Joshua, et al.
Veröffentlicht: (2025)
Improving Dictionary Learning with Gated Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
Survival of the Fittest Representation: A Case Study with Modular Addition
von: Ding, Xiaoman Delores, et al.
Veröffentlicht: (2024)
von: Ding, Xiaoman Delores, et al.
Veröffentlicht: (2024)
DafnyBench: A Benchmark for Formal Software Verification
von: Loughridge, Chloe, et al.
Veröffentlicht: (2024)
von: Loughridge, Chloe, et al.
Veröffentlicht: (2024)
Sparse Reconstruction of Wavefronts using an Over-Complete Phase Dictionary
von: Howard, S., et al.
Veröffentlicht: (2024)
von: Howard, S., et al.
Veröffentlicht: (2024)
Dense SAE Latents Are Features, Not Bugs
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
Towards Understanding Distilled Reasoning Models: A Representational Approach
von: Baek, David D., et al.
Veröffentlicht: (2025)
von: Baek, David D., et al.
Veröffentlicht: (2025)
Language Models Use Trigonometry to Do Addition
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
Language Models Represent Space and Time
von: Gurnee, Wes, et al.
Veröffentlicht: (2023)
von: Gurnee, Wes, et al.
Veröffentlicht: (2023)
Mirror Learning: A Unifying Framework of Policy Optimisation
von: Kuba, Jakub Grudzien, et al.
Veröffentlicht: (2022)
von: Kuba, Jakub Grudzien, et al.
Veröffentlicht: (2022)
A Neural Scaling Law from Lottery Ticket Ensembling
von: Liu, Ziming, et al.
Veröffentlicht: (2023)
von: Liu, Ziming, et al.
Veröffentlicht: (2023)
GenEFT: Understanding Statics and Dynamics of Model Generalization via Effective Theory
von: Baek, David D., et al.
Veröffentlicht: (2024)
von: Baek, David D., et al.
Veröffentlicht: (2024)
OptPDE: Discovering Novel Integrable Systems via AI-Human Collaboration
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2024)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2024)
Do Two AI Scientists Agree?
von: Fu, Xinghong, et al.
Veröffentlicht: (2025)
von: Fu, Xinghong, et al.
Veröffentlicht: (2025)
Bayesian Exploration Networks
von: Fellows, Mattie, et al.
Veröffentlicht: (2023)
von: Fellows, Mattie, et al.
Veröffentlicht: (2023)
Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models
von: Demircan, Can, et al.
Veröffentlicht: (2024)
von: Demircan, Can, et al.
Veröffentlicht: (2024)
Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
Investigating Representation Universality: Case Study on Genealogical Representations
von: Baek, David D., et al.
Veröffentlicht: (2024)
von: Baek, David D., et al.
Veröffentlicht: (2024)
Dictionary Learning: The Complexity of Learning Sparse Superposed Features with Feedback
von: Kumar, Akash
Veröffentlicht: (2025)
von: Kumar, Akash
Veröffentlicht: (2025)
Rethinking Out-of-Distribution Detection for Reinforcement Learning: Advancing Methods for Evaluation and Detection
von: Nasvytis, Linas, et al.
Veröffentlicht: (2024)
von: Nasvytis, Linas, et al.
Veröffentlicht: (2024)
Harmonic Loss Trains Interpretable AI Models
von: Baek, David D., et al.
Veröffentlicht: (2025)
von: Baek, David D., et al.
Veröffentlicht: (2025)
SwitchTab: Switched Autoencoders Are Effective Tabular Learners
von: Wu, Jing, et al.
Veröffentlicht: (2024)
von: Wu, Jing, et al.
Veröffentlicht: (2024)
DEEDEE: Fast and Scalable Out-of-Distribution Dynamics Detection
von: Aljaafari, Tala, et al.
Veröffentlicht: (2025)
von: Aljaafari, Tala, et al.
Veröffentlicht: (2025)
How Do Transformers "Do" Physics? Investigating the Simple Harmonic Oscillator
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2024)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2024)
Ensembling Sparse Autoencoders
von: Gadgil, Soham, et al.
Veröffentlicht: (2025)
von: Gadgil, Soham, et al.
Veröffentlicht: (2025)
Steering Language Model Refusal with Sparse Autoencoders
von: O'Brien, Kyle, et al.
Veröffentlicht: (2024)
von: O'Brien, Kyle, et al.
Veröffentlicht: (2024)
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning
von: Braun, Dan, et al.
Veröffentlicht: (2024)
von: Braun, Dan, et al.
Veröffentlicht: (2024)
Semi-Unified Sparse Dictionary Learning with Learnable Top-K LISTA and FISTA Encoders
von: Lin, Fengsheng, et al.
Veröffentlicht: (2025)
von: Lin, Fengsheng, et al.
Veröffentlicht: (2025)
Toward Identifiable Sparse Autoencoders
von: Nelson, Walter, et al.
Veröffentlicht: (2026)
von: Nelson, Walter, et al.
Veröffentlicht: (2026)
Analysis of Variational Sparse Autoencoders
von: Baker, Zachary, et al.
Veröffentlicht: (2025)
von: Baker, Zachary, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Decomposing The Dark Matter of Sparse Autoencoders
von: Engels, Joshua, et al.
Veröffentlicht: (2024) -
Low-Rank Adapting Models for Sparse Autoencoders
von: Chen, Matthew, et al.
Veröffentlicht: (2025) -
The Geometry of Concepts: Sparse Autoencoder Feature Structure
von: Li, Yuxiao, et al.
Veröffentlicht: (2024) -
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025) -
Not All Language Model Features Are One-Dimensionally Linear
von: Engels, Joshua, et al.
Veröffentlicht: (2024)