Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE)
Fuente:
arXiv
Saved in:
| Main Authors: | Brzozowski, Michał, Chung, Neo Christopher |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design
by: Brzozowski, Michał, et al.
Published: (2026)
by: Brzozowski, Michał, et al.
Published: (2026)
GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning
by: Mandica, Paolo, et al.
Published: (2026)
by: Mandica, Paolo, et al.
Published: (2026)
AlignSAE: Concept-Aligned Sparse Autoencoders
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing
by: Brzozowski, Michał, et al.
Published: (2026)
by: Brzozowski, Michał, et al.
Published: (2026)
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
by: Cao, Tue M., et al.
Published: (2026)
by: Cao, Tue M., et al.
Published: (2026)
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
by: Korznikov, Anton, et al.
Published: (2025)
by: Korznikov, Anton, et al.
Published: (2025)
SAE-FD: Sparse Autoencoder Feature Distillation for Continual Learning of Large Language Models
by: Zhang, Mingxu, et al.
Published: (2026)
by: Zhang, Mingxu, et al.
Published: (2026)
PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding
by: Koromilas, Panagiotis, et al.
Published: (2026)
by: Koromilas, Panagiotis, et al.
Published: (2026)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
by: Marks, Luke, et al.
Published: (2024)
by: Marks, Luke, et al.
Published: (2024)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
by: Cho, Seonglae, et al.
Published: (2025)
by: Cho, Seonglae, et al.
Published: (2025)
SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection
by: Zhang, Huopu, et al.
Published: (2025)
by: Zhang, Huopu, et al.
Published: (2025)
ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions
by: Poduval, Prathyush, et al.
Published: (2026)
by: Poduval, Prathyush, et al.
Published: (2026)
Sparse Autoencoders Trained on the Same Data Learn Different Features
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
SoftSAE: Dynamic Top-K Selection for Adaptive Sparse Autoencoders
by: Stępień, Jakub, et al.
Published: (2026)
by: Stępień, Jakub, et al.
Published: (2026)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
by: Chalnev, Sviatoslav, et al.
Published: (2024)
by: Chalnev, Sviatoslav, et al.
Published: (2024)
Regularizing Attention Scores with Bootstrapping
by: Chung, Neo Christopher, et al.
Published: (2026)
by: Chung, Neo Christopher, et al.
Published: (2026)
Improving Sparse Autoencoder with Dynamic Attention
by: Wang, Dongsheng, et al.
Published: (2026)
by: Wang, Dongsheng, et al.
Published: (2026)
Evolution of SAE Features Across Layers in LLMs
by: Balcells, Daniel, et al.
Published: (2024)
by: Balcells, Daniel, et al.
Published: (2024)
Sparse Autoencoder Features for Classifications and Transferability
by: Gallifant, Jack, et al.
Published: (2025)
by: Gallifant, Jack, et al.
Published: (2025)
Do Sparse Autoencoders Identify Reasoning Features in Language Models?
by: Ma, George, et al.
Published: (2026)
by: Ma, George, et al.
Published: (2026)
On the Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders
by: Simon, Elana, et al.
Published: (2026)
by: Simon, Elana, et al.
Published: (2026)
SplInterp: Improving our Understanding and Training of Sparse Autoencoders
by: Budd, Jeremy, et al.
Published: (2025)
by: Budd, Jeremy, et al.
Published: (2025)
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
by: Ayonrinde, Kola
Published: (2024)
by: Ayonrinde, Kola
Published: (2024)
Training Superior Sparse Autoencoders for Instruct Models
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
Improving Dictionary Learning with Gated Sparse Autoencoders
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
Data Whitening Improves Sparse Autoencoder Learning
by: Saraswatula, Ashwin, et al.
Published: (2025)
by: Saraswatula, Ashwin, et al.
Published: (2025)
MoRFI: Monotonic Sparse Autoencoder Feature Identification
by: Dimakopoulos, Dimitris, et al.
Published: (2026)
by: Dimakopoulos, Dimitris, et al.
Published: (2026)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
by: Bussmann, Bart, et al.
Published: (2025)
by: Bussmann, Bart, et al.
Published: (2025)
Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training
by: Li, T. Ed, et al.
Published: (2025)
by: Li, T. Ed, et al.
Published: (2025)
Feature Starvation as Geometric Instability in Sparse Autoencoders
by: Chaudhry, Faris, et al.
Published: (2026)
by: Chaudhry, Faris, et al.
Published: (2026)
The Geometry of Concepts: Sparse Autoencoder Feature Structure
by: Li, Yuxiao, et al.
Published: (2024)
by: Li, Yuxiao, et al.
Published: (2024)
RandAlign: A Parameter-Free Method for Regularizing Graph Convolutional Networks
by: Zhang, Haimin, et al.
Published: (2024)
by: Zhang, Haimin, et al.
Published: (2024)
GeoSAE: Geometric Prior-Guided Layer-Wise Sparse Autoencoder Annotation of Brain MRI Foundation Models
by: Nerrise, Favour, et al.
Published: (2026)
by: Nerrise, Favour, et al.
Published: (2026)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
by: Chanin, David, et al.
Published: (2025)
by: Chanin, David, et al.
Published: (2025)
Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution
by: Dubanowska, Zuzanna, et al.
Published: (2025)
by: Dubanowska, Zuzanna, et al.
Published: (2025)
Improving Robustness In Sparse Autoencoders via Masked Regularization
by: Narayanaswamy, Vivek, et al.
Published: (2026)
by: Narayanaswamy, Vivek, et al.
Published: (2026)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
by: Kurochkin, Vadim, et al.
Published: (2025)
by: Kurochkin, Vadim, et al.
Published: (2025)
Dense SAE Latents Are Features, Not Bugs
by: Sun, Xiaoqing, et al.
Published: (2025)
by: Sun, Xiaoqing, et al.
Published: (2025)
Class-Discriminative Attention Maps for Vision Transformers
by: Brocki, Lennart, et al.
Published: (2023)
by: Brocki, Lennart, et al.
Published: (2023)
Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
by: Yan, Xinyuan, et al.
Published: (2025)
by: Yan, Xinyuan, et al.
Published: (2025)
Similar Items
-
Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design
by: Brzozowski, Michał, et al.
Published: (2026) -
GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning
by: Mandica, Paolo, et al.
Published: (2026) -
AlignSAE: Concept-Aligned Sparse Autoencoders
by: Yang, Minglai, et al.
Published: (2025) -
Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing
by: Brzozowski, Michał, et al.
Published: (2026) -
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
by: Cao, Tue M., et al.
Published: (2026)