Evolution of SAE Features Across Layers in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Balcells, Daniel, Lerner, Benjamin, Oesterle, Michael, Ucar, Ediz, Heimersheim, Stefan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
von: Fillingham, Sean P., et al.
Veröffentlicht: (2025)
von: Fillingham, Sean P., et al.
Veröffentlicht: (2025)
You can remove GPT2's LayerNorm by fine-tuning
von: Heimersheim, Stefan
Veröffentlicht: (2024)
von: Heimersheim, Stefan
Veröffentlicht: (2024)
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
von: Lee, Daniel J., et al.
Veröffentlicht: (2024)
von: Lee, Daniel J., et al.
Veröffentlicht: (2024)
How to use and interpret activation patching
von: Heimersheim, Stefan, et al.
Veröffentlicht: (2024)
von: Heimersheim, Stefan, et al.
Veröffentlicht: (2024)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
von: Baroni, Luca, et al.
Veröffentlicht: (2025)
von: Baroni, Luca, et al.
Veröffentlicht: (2025)
Characterizing stable regions in the residual stream of LLMs
von: Janiak, Jett, et al.
Veröffentlicht: (2024)
von: Janiak, Jett, et al.
Veröffentlicht: (2024)
Calibration Across Layers: Understanding Calibration Evolution in LLMs
von: Joshi, Abhinav, et al.
Veröffentlicht: (2025)
von: Joshi, Abhinav, et al.
Veröffentlicht: (2025)
Detecting Strategic Deception Using Linear Probes
von: Goldowsky-Dill, Nicholas, et al.
Veröffentlicht: (2025)
von: Goldowsky-Dill, Nicholas, et al.
Veröffentlicht: (2025)
Mechanistic Permutability: Match Features Across Layers
von: Balagansky, Nikita, et al.
Veröffentlicht: (2024)
von: Balagansky, Nikita, et al.
Veröffentlicht: (2024)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026)
Dense SAE Latents Are Features, Not Bugs
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
Tokenized SAEs: Disentangling SAE Reconstructions
von: Dooms, Thomas, et al.
Veröffentlicht: (2025)
von: Dooms, Thomas, et al.
Veröffentlicht: (2025)
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
von: Braun, Dan, et al.
Veröffentlicht: (2025)
von: Braun, Dan, et al.
Veröffentlicht: (2025)
Compositional Literary Primitives in Instruction-Tuned LLMs: Cross-Architectural SAE Features for Self, Style, and Affect
von: Presa, Joao Paulo Cavalcante, et al.
Veröffentlicht: (2026)
von: Presa, Joao Paulo Cavalcante, et al.
Veröffentlicht: (2026)
ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions
von: Poduval, Prathyush, et al.
Veröffentlicht: (2026)
von: Poduval, Prathyush, et al.
Veröffentlicht: (2026)
SAE-FD: Sparse Autoencoder Feature Distillation for Continual Learning of Large Language Models
von: Zhang, Mingxu, et al.
Veröffentlicht: (2026)
von: Zhang, Mingxu, et al.
Veröffentlicht: (2026)
The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
Benchmarking Deception Probes via Black-to-White Performance Boosts
von: Parrack, Avi, et al.
Veröffentlicht: (2025)
von: Parrack, Avi, et al.
Veröffentlicht: (2025)
Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts
von: Bhargav, Samaksh, et al.
Veröffentlicht: (2025)
von: Bhargav, Samaksh, et al.
Veröffentlicht: (2025)
PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding
von: Koromilas, Panagiotis, et al.
Veröffentlicht: (2026)
von: Koromilas, Panagiotis, et al.
Veröffentlicht: (2026)
Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE)
von: Brzozowski, Michał, et al.
Veröffentlicht: (2026)
von: Brzozowski, Michał, et al.
Veröffentlicht: (2026)
Evaluating SAE interpretability without explanations
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
SAE: Single Architecture Ensemble Neural Networks
von: Ferianc, Martin, et al.
Veröffentlicht: (2024)
von: Ferianc, Martin, et al.
Veröffentlicht: (2024)
Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs
von: Jiang, Jingzhou, et al.
Veröffentlicht: (2026)
von: Jiang, Jingzhou, et al.
Veröffentlicht: (2026)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
Concept-SAE: Active Causal Probing of Visual Model Behavior
von: Ding, Jianrong, et al.
Veröffentlicht: (2025)
von: Ding, Jianrong, et al.
Veröffentlicht: (2025)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection
von: Zhang, Huopu, et al.
Veröffentlicht: (2025)
von: Zhang, Huopu, et al.
Veröffentlicht: (2025)
AlignSAE: Concept-Aligned Sparse Autoencoders
von: Yang, Minglai, et al.
Veröffentlicht: (2025)
von: Yang, Minglai, et al.
Veröffentlicht: (2025)
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024)
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026)
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026)
The Optimization Landscape of SGD Across the Feature Learning Strength
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds
von: Wu, Honghan, et al.
Veröffentlicht: (2026)
von: Wu, Honghan, et al.
Veröffentlicht: (2026)
TimeSAE: Sparse Decoding for Faithful Explanations of Black-Box Time Series Models
von: Oublal, Khalid, et al.
Veröffentlicht: (2026)
von: Oublal, Khalid, et al.
Veröffentlicht: (2026)
Not Only the Last-Layer Features for Spurious Correlations: All Layer Deep Feature Reweighting
von: Hameed, Humza Wajid, et al.
Veröffentlicht: (2024)
von: Hameed, Humza Wajid, et al.
Veröffentlicht: (2024)
GeoSAE: Geometric Prior-Guided Layer-Wise Sparse Autoencoder Annotation of Brain MRI Foundation Models
von: Nerrise, Favour, et al.
Veröffentlicht: (2026)
von: Nerrise, Favour, et al.
Veröffentlicht: (2026)
Deep Reinforcement Learning for Local Path Following of an Autonomous Formula SAE Vehicle
von: Merton, Harvey, et al.
Veröffentlicht: (2024)
von: Merton, Harvey, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
von: Fillingham, Sean P., et al.
Veröffentlicht: (2025) -
You can remove GPT2's LayerNorm by fine-tuning
von: Heimersheim, Stefan
Veröffentlicht: (2024) -
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024) -
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
von: Lee, Daniel J., et al.
Veröffentlicht: (2024) -
How to use and interpret activation patching
von: Heimersheim, Stefan, et al.
Veröffentlicht: (2024)