Tokenized SAEs: Disentangling SAE Reconstructions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dooms, Thomas, Wilhelm, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024)
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024)
Finding Manifolds With Bilinear Autoencoders
von: Dooms, Thomas, et al.
Veröffentlicht: (2025)
von: Dooms, Thomas, et al.
Veröffentlicht: (2025)
Distribution-Aware Feature Selection for SAEs
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
The Rate-Distortion-Polysemanticity Tradeoff in SAEs
von: Mencattini, Tommaso, et al.
Veröffentlicht: (2026)
von: Mencattini, Tommaso, et al.
Veröffentlicht: (2026)
Analyzing (In)Abilities of SAEs via Formal Languages
von: Menon, Abhinav, et al.
Veröffentlicht: (2024)
von: Menon, Abhinav, et al.
Veröffentlicht: (2024)
Fractals made Practical: Denoising Diffusion as Partitioned Iterated Function Systems
von: Dooms, Ann
Veröffentlicht: (2026)
von: Dooms, Ann
Veröffentlicht: (2026)
Weight-based Decomposition: A Case for Bilinear MLPs
von: Pearce, Michael T., et al.
Veröffentlicht: (2024)
von: Pearce, Michael T., et al.
Veröffentlicht: (2024)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
von: Lee, Daniel J., et al.
Veröffentlicht: (2024)
von: Lee, Daniel J., et al.
Veröffentlicht: (2024)
Bilinear autoencoders find interpretable manifolds
von: Dooms, Thomas, et al.
Veröffentlicht: (2026)
von: Dooms, Thomas, et al.
Veröffentlicht: (2026)
Residual Stream Analysis with Multi-Layer SAEs
von: Lawson, Tim, et al.
Veröffentlicht: (2024)
von: Lawson, Tim, et al.
Veröffentlicht: (2024)
Compositionality Unlocks Deep Interpretable Models
von: Dooms, Thomas, et al.
Veröffentlicht: (2025)
von: Dooms, Thomas, et al.
Veröffentlicht: (2025)
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
von: Korznikov, Anton, et al.
Veröffentlicht: (2026)
von: Korznikov, Anton, et al.
Veröffentlicht: (2026)
Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design
von: Brzozowski, Michał, et al.
Veröffentlicht: (2026)
von: Brzozowski, Michał, et al.
Veröffentlicht: (2026)
Resa: Transparent Reasoning Models via SAEs
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)
Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
von: Ahsan, Hiba, et al.
Veröffentlicht: (2025)
von: Ahsan, Hiba, et al.
Veröffentlicht: (2025)
Features Emerge as Discrete States: The First Application of SAEs to 3D Representations
von: Miao, Albert, et al.
Veröffentlicht: (2025)
von: Miao, Albert, et al.
Veröffentlicht: (2025)
SAEs Are Good for Steering -- If You Select the Right Features
von: Arad, Dana, et al.
Veröffentlicht: (2025)
von: Arad, Dana, et al.
Veröffentlicht: (2025)
Teach Old SAEs New Domain Tricks with Boosting
von: Koriagin, Nikita, et al.
Veröffentlicht: (2025)
von: Koriagin, Nikita, et al.
Veröffentlicht: (2025)
From Mechanistic to Compositional Interpretability
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
Bilinear MLPs enable weight-based mechanistic interpretability
von: Pearce, Michael T., et al.
Veröffentlicht: (2024)
von: Pearce, Michael T., et al.
Veröffentlicht: (2024)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
von: Song, Xiangchen, et al.
Veröffentlicht: (2025)
von: Song, Xiangchen, et al.
Veröffentlicht: (2025)
Evolution of SAE Features Across Layers in LLMs
von: Balcells, Daniel, et al.
Veröffentlicht: (2024)
von: Balcells, Daniel, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability with SAEs: Probing Religion, Violence, and Geography in Large Language Models
von: Simbeck, Katharina, et al.
Veröffentlicht: (2025)
von: Simbeck, Katharina, et al.
Veröffentlicht: (2025)
Evaluating SAE interpretability without explanations
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
SAE: Single Architecture Ensemble Neural Networks
von: Ferianc, Martin, et al.
Veröffentlicht: (2024)
von: Ferianc, Martin, et al.
Veröffentlicht: (2024)
Behavior Tokens Speak Louder: Disentangled Explainable Recommendation with Behavior Vocabulary
von: Feng, Xinshun, et al.
Veröffentlicht: (2025)
von: Feng, Xinshun, et al.
Veröffentlicht: (2025)
When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
von: Gonzalez, ML Nissen, et al.
Veröffentlicht: (2026)
von: Gonzalez, ML Nissen, et al.
Veröffentlicht: (2026)
Deep Reinforcement Learning for Local Path Following of an Autonomous Formula SAE Vehicle
von: Merton, Harvey, et al.
Veröffentlicht: (2024)
von: Merton, Harvey, et al.
Veröffentlicht: (2024)
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
von: Fillingham, Sean P., et al.
Veröffentlicht: (2025)
von: Fillingham, Sean P., et al.
Veröffentlicht: (2025)
MetaSAEs: Joint Training with a Decomposability Penalty Produces More Atomic Sparse Autoencoder Latents
von: Levinson, Matthew
Veröffentlicht: (2026)
von: Levinson, Matthew
Veröffentlicht: (2026)
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
Concept-SAE: Active Causal Probing of Visual Model Behavior
von: Ding, Jianrong, et al.
Veröffentlicht: (2025)
von: Ding, Jianrong, et al.
Veröffentlicht: (2025)
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
AlignSAE: Concept-Aligned Sparse Autoencoders
von: Yang, Minglai, et al.
Veröffentlicht: (2025)
von: Yang, Minglai, et al.
Veröffentlicht: (2025)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
Dense SAE Latents Are Features, Not Bugs
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds
von: Wu, Honghan, et al.
Veröffentlicht: (2026)
von: Wu, Honghan, et al.
Veröffentlicht: (2026)
TimeSAE: Sparse Decoding for Faithful Explanations of Black-Box Time Series Models
von: Oublal, Khalid, et al.
Veröffentlicht: (2026)
von: Oublal, Khalid, et al.
Veröffentlicht: (2026)
SAE-FD: Sparse Autoencoder Feature Distillation for Continual Learning of Large Language Models
von: Zhang, Mingxu, et al.
Veröffentlicht: (2026)
von: Zhang, Mingxu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024) -
Finding Manifolds With Bilinear Autoencoders
von: Dooms, Thomas, et al.
Veröffentlicht: (2025) -
Distribution-Aware Feature Selection for SAEs
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025) -
The Rate-Distortion-Polysemanticity Tradeoff in SAEs
von: Mencattini, Tommaso, et al.
Veröffentlicht: (2026) -
Analyzing (In)Abilities of SAEs via Formal Languages
von: Menon, Abhinav, et al.
Veröffentlicht: (2024)