Sparse Autoencoders Trained on the Same Data Learn Different Features
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Paulo, Gonçalo, Belrose, Nora |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transcoders Beat Sparse Autoencoders for Interpretability
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
Evaluating SAE interpretability without explanations
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
Partially Rewriting a Transformer in Natural Language
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
Automatically Interpreting Millions of Features in Large Language Models
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
Does Transformer Interpretability Transfer to RNNs?
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
Slowing Learning by Erasing Simple Features
von: Quirke, Lucia, et al.
Veröffentlicht: (2025)
von: Quirke, Lucia, et al.
Veröffentlicht: (2025)
Understanding Gradient Descent through the Training Jacobian
von: Belrose, Nora, et al.
Veröffentlicht: (2024)
von: Belrose, Nora, et al.
Veröffentlicht: (2024)
Estimating the Probability of Sampling a Trained Neural Network at Random
von: Scherlis, Adam, et al.
Veröffentlicht: (2025)
von: Scherlis, Adam, et al.
Veröffentlicht: (2025)
Binary Sparse Coding for Interpretability
von: Quirke, Lucia, et al.
Veröffentlicht: (2025)
von: Quirke, Lucia, et al.
Veröffentlicht: (2025)
Converting MLPs into Polynomials in Closed Form
von: Belrose, Nora, et al.
Veröffentlicht: (2025)
von: Belrose, Nora, et al.
Veröffentlicht: (2025)
Balancing Label Quantity and Quality for Scalable Elicitation
von: Mallen, Alex, et al.
Veröffentlicht: (2024)
von: Mallen, Alex, et al.
Veröffentlicht: (2024)
Examining Two Hop Reasoning Through Information Content Scaling
von: Johnston, David, et al.
Veröffentlicht: (2025)
von: Johnston, David, et al.
Veröffentlicht: (2025)
Refusal in LLMs is an Affine Function
von: Marshall, Thomas, et al.
Veröffentlicht: (2024)
von: Marshall, Thomas, et al.
Veröffentlicht: (2024)
Neural Networks Learn Statistics of Increasing Complexity
von: Belrose, Nora, et al.
Veröffentlicht: (2024)
von: Belrose, Nora, et al.
Veröffentlicht: (2024)
Mechanistic Anomaly Detection for "Quirky" Language Models
von: Johnston, David O., et al.
Veröffentlicht: (2025)
von: Johnston, David O., et al.
Veröffentlicht: (2025)
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
von: Patel, Het, et al.
Veröffentlicht: (2026)
von: Patel, Het, et al.
Veröffentlicht: (2026)
Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models
von: Demircan, Can, et al.
Veröffentlicht: (2024)
von: Demircan, Can, et al.
Veröffentlicht: (2024)
Grokking vs. Learning: Same Features, Different Encodings
von: Manning-Coe, Dmitry, et al.
Veröffentlicht: (2025)
von: Manning-Coe, Dmitry, et al.
Veröffentlicht: (2025)
Data Whitening Improves Sparse Autoencoder Learning
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025)
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025)
Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
von: Paek, Nathan, et al.
Veröffentlicht: (2025)
von: Paek, Nathan, et al.
Veröffentlicht: (2025)
Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE)
von: Brzozowski, Michał, et al.
Veröffentlicht: (2026)
von: Brzozowski, Michał, et al.
Veröffentlicht: (2026)
Sparse Autoencoder Features for Classifications and Transferability
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
Eliciting Latent Knowledge from Quirky Language Models
von: Mallen, Alex, et al.
Veröffentlicht: (2023)
von: Mallen, Alex, et al.
Veröffentlicht: (2023)
SAE-FD: Sparse Autoencoder Feature Distillation for Continual Learning of Large Language Models
von: Zhang, Mingxu, et al.
Veröffentlicht: (2026)
von: Zhang, Mingxu, et al.
Veröffentlicht: (2026)
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
Do Sparse Autoencoders Identify Reasoning Features in Language Models?
von: Ma, George, et al.
Veröffentlicht: (2026)
von: Ma, George, et al.
Veröffentlicht: (2026)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
von: Marks, Luke, et al.
Veröffentlicht: (2024)
von: Marks, Luke, et al.
Veröffentlicht: (2024)
On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders
von: Simon, Elana, et al.
Veröffentlicht: (2026)
von: Simon, Elana, et al.
Veröffentlicht: (2026)
Training Superior Sparse Autoencoders for Instruct Models
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
von: Ayonrinde, Kola
Veröffentlicht: (2024)
von: Ayonrinde, Kola
Veröffentlicht: (2024)
MoRFI: Monotonic Sparse Autoencoder Feature Identification
von: Dimakopoulos, Dimitris, et al.
Veröffentlicht: (2026)
von: Dimakopoulos, Dimitris, et al.
Veröffentlicht: (2026)
Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training
von: Li, T. Ed, et al.
Veröffentlicht: (2025)
von: Li, T. Ed, et al.
Veröffentlicht: (2025)
Feature Starvation as Geometric Instability in Sparse Autoencoders
von: Chaudhry, Faris, et al.
Veröffentlicht: (2026)
von: Chaudhry, Faris, et al.
Veröffentlicht: (2026)
The Geometry of Concepts: Sparse Autoencoder Feature Structure
von: Li, Yuxiao, et al.
Veröffentlicht: (2024)
von: Li, Yuxiao, et al.
Veröffentlicht: (2024)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025)
von: Chanin, David, et al.
Veröffentlicht: (2025)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
Efficient Dictionary Learning with Switch Sparse Autoencoders
von: Mudide, Anish, et al.
Veröffentlicht: (2024)
von: Mudide, Anish, et al.
Veröffentlicht: (2024)
Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
von: Yan, Xinyuan, et al.
Veröffentlicht: (2025)
von: Yan, Xinyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Transcoders Beat Sparse Autoencoders for Interpretability
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025) -
Evaluating SAE interpretability without explanations
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025) -
Partially Rewriting a Transformer in Natural Language
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025) -
Automatically Interpreting Millions of Features in Large Language Models
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024) -
Does Transformer Interpretability Transfer to RNNs?
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)