Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design
Fuente:
arXiv
Saved in:
| Main Authors: | Brzozowski, Michał, Chung, Neo Christopher |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE)
by: Brzozowski, Michał, et al.
Published: (2026)
by: Brzozowski, Michał, et al.
Published: (2026)
Archetypal Analysis++: Rethinking the Initialization Strategy
by: Mair, Sebastian, et al.
Published: (2023)
by: Mair, Sebastian, et al.
Published: (2023)
Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing
by: Brzozowski, Michał, et al.
Published: (2026)
by: Brzozowski, Michał, et al.
Published: (2026)
GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning
by: Mandica, Paolo, et al.
Published: (2026)
by: Mandica, Paolo, et al.
Published: (2026)
A Survey on Archetypal Analysis
by: Alcacer, Aleix, et al.
Published: (2025)
by: Alcacer, Aleix, et al.
Published: (2025)
Archetypal Analysis for Binary Data
by: Wedenborg, A. Emilie J., et al.
Published: (2025)
by: Wedenborg, A. Emilie J., et al.
Published: (2025)
Incorporating Fairness Constraints into Archetypal Analysis
by: Alcacer, Aleix, et al.
Published: (2025)
by: Alcacer, Aleix, et al.
Published: (2025)
Modeling Human Responses by Ordinal Archetypal Analysis
by: Wedenborg, Anna Emilie J., et al.
Published: (2024)
by: Wedenborg, Anna Emilie J., et al.
Published: (2024)
Sparse NMF with Archetypal Regularization: Computational and Robustness Properties
by: Behdin, Kayhan, et al.
Published: (2021)
by: Behdin, Kayhan, et al.
Published: (2021)
Archetypal cases for questionnaires with nominal multiple choice questions
by: Alcacer, Aleix, et al.
Published: (2026)
by: Alcacer, Aleix, et al.
Published: (2026)
Level the Level: Balancing Game Levels for Asymmetric Player Archetypes With Reinforcement Learning
by: Rupp, Florian, et al.
Published: (2025)
by: Rupp, Florian, et al.
Published: (2025)
Archetypal Graph Generative Models: Explainable and Identifiable Communities via Anchor-Dominant Convex Hulls
by: Nakis, Nikolaos, et al.
Published: (2026)
by: Nakis, Nikolaos, et al.
Published: (2026)
Lattice: A Confidence-Gated Hybrid System for Uncertainty-Aware Sequential Prediction with Behavioral Archetypes
by: Bannis, Lorian
Published: (2026)
by: Bannis, Lorian
Published: (2026)
Discovering EV Charging Site Archetypes Through Few Shot Forecasting: The First U.S.-Wide Study
by: Nikhal, Kshitij, et al.
Published: (2025)
by: Nikhal, Kshitij, et al.
Published: (2025)
Unveiling Gamer Archetypes through Multi modal feature Correlations and Unsupervised Learning
by: Kanwal, Moona, et al.
Published: (2025)
by: Kanwal, Moona, et al.
Published: (2025)
OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation
by: Yang, Haochen, et al.
Published: (2026)
by: Yang, Haochen, et al.
Published: (2026)
Robust Player-Conditional Champion Ranking for League of Legends: Style Similarity, Mastery Priors, and Archetype-Constrained Discovery
by: Heo, Min, et al.
Published: (2026)
by: Heo, Min, et al.
Published: (2026)
Riemannian Archetypal Analysis: Interpretable non-linear data analysis on deformed star distributions
by: Diepeveen, Willem, et al.
Published: (2026)
by: Diepeveen, Willem, et al.
Published: (2026)
Structural Pass Analysis in Football: Learning Pass Archetypes and Tactical Impact from Spatio-Temporal Tracking Data
by: Karakuş, Oktay, et al.
Published: (2026)
by: Karakuş, Oktay, et al.
Published: (2026)
The Rate-Distortion-Polysemanticity Tradeoff in SAEs
by: Mencattini, Tommaso, et al.
Published: (2026)
by: Mencattini, Tommaso, et al.
Published: (2026)
Tokenized SAEs: Disentangling SAE Reconstructions
by: Dooms, Thomas, et al.
Published: (2025)
by: Dooms, Thomas, et al.
Published: (2025)
Distribution-Aware Feature Selection for SAEs
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Analyzing (In)Abilities of SAEs via Formal Languages
by: Menon, Abhinav, et al.
Published: (2024)
by: Menon, Abhinav, et al.
Published: (2024)
Regularizing Attention Scores with Bootstrapping
by: Chung, Neo Christopher, et al.
Published: (2026)
by: Chung, Neo Christopher, et al.
Published: (2026)
Self-Ablating Transformers: More Interpretability, Less Sparsity
by: Ferrao, Jeremias, et al.
Published: (2025)
by: Ferrao, Jeremias, et al.
Published: (2025)
Mother Archetype
by: Ganganmale, Pramod Akaram
Published: (2026)
by: Ganganmale, Pramod Akaram
Published: (2026)
Residual Stream Analysis with Multi-Layer SAEs
by: Lawson, Tim, et al.
Published: (2024)
by: Lawson, Tim, et al.
Published: (2024)
Bayesian algorithmic perfumery: A Hierarchical Relevance Vector Machine for the Estimation of Personalized Fragrance Preferences based on Three Sensory Layers and Jungian Personality Archetypes
by: Martinez, Rolando Gonzales
Published: (2024)
by: Martinez, Rolando Gonzales
Published: (2024)
Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution
by: Dubanowska, Zuzanna, et al.
Published: (2025)
by: Dubanowska, Zuzanna, et al.
Published: (2025)
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
by: Korznikov, Anton, et al.
Published: (2026)
by: Korznikov, Anton, et al.
Published: (2026)
Class-Discriminative Attention Maps for Vision Transformers
by: Brocki, Lennart, et al.
Published: (2023)
by: Brocki, Lennart, et al.
Published: (2023)
Resa: Transparent Reasoning Models via SAEs
by: Wang, Shangshang, et al.
Published: (2025)
by: Wang, Shangshang, et al.
Published: (2025)
Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
by: Ahsan, Hiba, et al.
Published: (2025)
by: Ahsan, Hiba, et al.
Published: (2025)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
by: Lee, Daniel J., et al.
Published: (2024)
by: Lee, Daniel J., et al.
Published: (2024)
Features Emerge as Discrete States: The First Application of SAEs to 3D Representations
by: Miao, Albert, et al.
Published: (2025)
by: Miao, Albert, et al.
Published: (2025)
TumorArchetypeR
by: Lütge, Mechthild, et al.
Published: (2026)
by: Lütge, Mechthild, et al.
Published: (2026)
SAEs Are Good for Steering -- If You Select the Right Features
by: Arad, Dana, et al.
Published: (2025)
by: Arad, Dana, et al.
Published: (2025)
Teach Old SAEs New Domain Tricks with Boosting
by: Koriagin, Nikita, et al.
Published: (2025)
by: Koriagin, Nikita, et al.
Published: (2025)
What Sort of a Thing is an Archetype? Archetypes, Complexes and Self‐Organization Revisited
by: Patricia Skar
Published: (2025)
by: Patricia Skar
Published: (2025)
Ablate and Rescue: A Causal Analysis of Residual Stream Hyper-Connections
by: Peng, William, et al.
Published: (2026)
by: Peng, William, et al.
Published: (2026)
Similar Items
-
Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE)
by: Brzozowski, Michał, et al.
Published: (2026) -
Archetypal Analysis++: Rethinking the Initialization Strategy
by: Mair, Sebastian, et al.
Published: (2023) -
Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing
by: Brzozowski, Michał, et al.
Published: (2026) -
GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning
by: Mandica, Paolo, et al.
Published: (2026) -
A Survey on Archetypal Analysis
by: Alcacer, Aleix, et al.
Published: (2025)