Tokenized SAEs: Disentangling SAE Reconstructions
Fuente:
arXiv
Guardado en:
| Autores principales: | Dooms, Thomas, Wilhelm, Daniel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
por: Ayonrinde, Kola, et al.
Publicado: (2024)
por: Ayonrinde, Kola, et al.
Publicado: (2024)
Finding Manifolds With Bilinear Autoencoders
por: Dooms, Thomas, et al.
Publicado: (2025)
por: Dooms, Thomas, et al.
Publicado: (2025)
Distribution-Aware Feature Selection for SAEs
por: Oozeer, Narmeen, et al.
Publicado: (2025)
por: Oozeer, Narmeen, et al.
Publicado: (2025)
The Rate-Distortion-Polysemanticity Tradeoff in SAEs
por: Mencattini, Tommaso, et al.
Publicado: (2026)
por: Mencattini, Tommaso, et al.
Publicado: (2026)
Analyzing (In)Abilities of SAEs via Formal Languages
por: Menon, Abhinav, et al.
Publicado: (2024)
por: Menon, Abhinav, et al.
Publicado: (2024)
Fractals made Practical: Denoising Diffusion as Partitioned Iterated Function Systems
por: Dooms, Ann
Publicado: (2026)
por: Dooms, Ann
Publicado: (2026)
Weight-based Decomposition: A Case for Bilinear MLPs
por: Pearce, Michael T., et al.
Publicado: (2024)
por: Pearce, Michael T., et al.
Publicado: (2024)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
por: Lee, Daniel J., et al.
Publicado: (2024)
por: Lee, Daniel J., et al.
Publicado: (2024)
Bilinear autoencoders find interpretable manifolds
por: Dooms, Thomas, et al.
Publicado: (2026)
por: Dooms, Thomas, et al.
Publicado: (2026)
Residual Stream Analysis with Multi-Layer SAEs
por: Lawson, Tim, et al.
Publicado: (2024)
por: Lawson, Tim, et al.
Publicado: (2024)
Compositionality Unlocks Deep Interpretable Models
por: Dooms, Thomas, et al.
Publicado: (2025)
por: Dooms, Thomas, et al.
Publicado: (2025)
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
por: Korznikov, Anton, et al.
Publicado: (2026)
por: Korznikov, Anton, et al.
Publicado: (2026)
Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design
por: Brzozowski, Michał, et al.
Publicado: (2026)
por: Brzozowski, Michał, et al.
Publicado: (2026)
Resa: Transparent Reasoning Models via SAEs
por: Wang, Shangshang, et al.
Publicado: (2025)
por: Wang, Shangshang, et al.
Publicado: (2025)
Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
por: Ahsan, Hiba, et al.
Publicado: (2025)
por: Ahsan, Hiba, et al.
Publicado: (2025)
Features Emerge as Discrete States: The First Application of SAEs to 3D Representations
por: Miao, Albert, et al.
Publicado: (2025)
por: Miao, Albert, et al.
Publicado: (2025)
SAEs Are Good for Steering -- If You Select the Right Features
por: Arad, Dana, et al.
Publicado: (2025)
por: Arad, Dana, et al.
Publicado: (2025)
Teach Old SAEs New Domain Tricks with Boosting
por: Koriagin, Nikita, et al.
Publicado: (2025)
por: Koriagin, Nikita, et al.
Publicado: (2025)
From Mechanistic to Compositional Interpretability
por: Gauderis, Ward, et al.
Publicado: (2026)
por: Gauderis, Ward, et al.
Publicado: (2026)
Bilinear MLPs enable weight-based mechanistic interpretability
por: Pearce, Michael T., et al.
Publicado: (2024)
por: Pearce, Michael T., et al.
Publicado: (2024)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
por: Song, Xiangchen, et al.
Publicado: (2025)
por: Song, Xiangchen, et al.
Publicado: (2025)
Evolution of SAE Features Across Layers in LLMs
por: Balcells, Daniel, et al.
Publicado: (2024)
por: Balcells, Daniel, et al.
Publicado: (2024)
Mechanistic Interpretability with SAEs: Probing Religion, Violence, and Geography in Large Language Models
por: Simbeck, Katharina, et al.
Publicado: (2025)
por: Simbeck, Katharina, et al.
Publicado: (2025)
Evaluating SAE interpretability without explanations
por: Paulo, Gonçalo, et al.
Publicado: (2025)
por: Paulo, Gonçalo, et al.
Publicado: (2025)
SAE: Single Architecture Ensemble Neural Networks
por: Ferianc, Martin, et al.
Publicado: (2024)
por: Ferianc, Martin, et al.
Publicado: (2024)
Behavior Tokens Speak Louder: Disentangled Explainable Recommendation with Behavior Vocabulary
por: Feng, Xinshun, et al.
Publicado: (2025)
por: Feng, Xinshun, et al.
Publicado: (2025)
When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
por: Gonzalez, ML Nissen, et al.
Publicado: (2026)
por: Gonzalez, ML Nissen, et al.
Publicado: (2026)
Deep Reinforcement Learning for Local Path Following of an Autonomous Formula SAE Vehicle
por: Merton, Harvey, et al.
Publicado: (2024)
por: Merton, Harvey, et al.
Publicado: (2024)
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
por: Fillingham, Sean P., et al.
Publicado: (2025)
por: Fillingham, Sean P., et al.
Publicado: (2025)
MetaSAEs: Joint Training with a Decomposability Penalty Produces More Atomic Sparse Autoencoder Latents
por: Levinson, Matthew
Publicado: (2026)
por: Levinson, Matthew
Publicado: (2026)
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
por: Korznikov, Anton, et al.
Publicado: (2025)
por: Korznikov, Anton, et al.
Publicado: (2025)
Concept-SAE: Active Causal Probing of Visual Model Behavior
por: Ding, Jianrong, et al.
Publicado: (2025)
por: Ding, Jianrong, et al.
Publicado: (2025)
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
por: Giglemiani, Giorgi, et al.
Publicado: (2024)
por: Giglemiani, Giorgi, et al.
Publicado: (2024)
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
por: Cao, Tue M., et al.
Publicado: (2026)
por: Cao, Tue M., et al.
Publicado: (2026)
AlignSAE: Concept-Aligned Sparse Autoencoders
por: Yang, Minglai, et al.
Publicado: (2025)
por: Yang, Minglai, et al.
Publicado: (2025)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
por: Muhamed, Aashiq, et al.
Publicado: (2025)
por: Muhamed, Aashiq, et al.
Publicado: (2025)
Dense SAE Latents Are Features, Not Bugs
por: Sun, Xiaoqing, et al.
Publicado: (2025)
por: Sun, Xiaoqing, et al.
Publicado: (2025)
HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds
por: Wu, Honghan, et al.
Publicado: (2026)
por: Wu, Honghan, et al.
Publicado: (2026)
TimeSAE: Sparse Decoding for Faithful Explanations of Black-Box Time Series Models
por: Oublal, Khalid, et al.
Publicado: (2026)
por: Oublal, Khalid, et al.
Publicado: (2026)
SAE-FD: Sparse Autoencoder Feature Distillation for Continual Learning of Large Language Models
por: Zhang, Mingxu, et al.
Publicado: (2026)
por: Zhang, Mingxu, et al.
Publicado: (2026)
Ejemplares similares
-
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
por: Ayonrinde, Kola, et al.
Publicado: (2024) -
Finding Manifolds With Bilinear Autoencoders
por: Dooms, Thomas, et al.
Publicado: (2025) -
Distribution-Aware Feature Selection for SAEs
por: Oozeer, Narmeen, et al.
Publicado: (2025) -
The Rate-Distortion-Polysemanticity Tradeoff in SAEs
por: Mencattini, Tommaso, et al.
Publicado: (2026) -
Analyzing (In)Abilities of SAEs via Formal Languages
por: Menon, Abhinav, et al.
Publicado: (2024)