Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
Fuente:
arXiv
Saved in:
| Main Authors: | Korznikov, Anton, Galichin, Andrey, Dontsov, Alexey, Rogov, Oleg, Oseledets, Ivan, Tutubalina, Elena |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
by: Korznikov, Anton, et al.
Published: (2025)
by: Korznikov, Anton, et al.
Published: (2025)
The Rogue Scalpel: Activation Steering Compromises LLM Safety
by: Korznikov, Anton, et al.
Published: (2025)
by: Korznikov, Anton, et al.
Published: (2025)
I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
by: Galichin, Andrey, et al.
Published: (2025)
by: Galichin, Andrey, et al.
Published: (2025)
GLiRA: Black-Box Membership Inference Attack via Knowledge Distillation
by: Galichin, Andrey V., et al.
Published: (2024)
by: Galichin, Andrey V., et al.
Published: (2024)
Spread them Apart: Towards Robust Watermarking of Generated Content
by: Pautov, Mikhail, et al.
Published: (2025)
by: Pautov, Mikhail, et al.
Published: (2025)
CLEAR: Character Unlearning in Textual and Visual Modalities
by: Dontsov, Alexey, et al.
Published: (2024)
by: Dontsov, Alexey, et al.
Published: (2024)
Sanity Checks for Explanation Uncertainty
by: Valdenegro-Toro, Matias, et al.
Published: (2024)
by: Valdenegro-Toro, Matias, et al.
Published: (2024)
GigaEvo: An Open Source Optimization Framework Powered By LLMs And Evolution Algorithms
by: Khrulkov, Valentin, et al.
Published: (2025)
by: Khrulkov, Valentin, et al.
Published: (2025)
Sanity Checks for Agentic Data Science
by: Rewolinski, Zachary T., et al.
Published: (2026)
by: Rewolinski, Zachary T., et al.
Published: (2026)
Transcoders Beat Sparse Autoencoders for Interpretability
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Confidence Estimation for Error Detection in Text-to-SQL Systems
by: Somov, Oleg, et al.
Published: (2025)
by: Somov, Oleg, et al.
Published: (2025)
Evaluating Neuron Explanations: A Unified Framework with Sanity Checks
by: Oikarinen, Tuomas, et al.
Published: (2025)
by: Oikarinen, Tuomas, et al.
Published: (2025)
A Fresh Look at Sanity Checks for Saliency Maps
by: Hedström, Anna, et al.
Published: (2024)
by: Hedström, Anna, et al.
Published: (2024)
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
by: Lee, Daniel J., et al.
Published: (2024)
by: Lee, Daniel J., et al.
Published: (2024)
MetaSAEs: Joint Training with a Decomposability Penalty Produces More Atomic Sparse Autoencoder Latents
by: Levinson, Matthew
Published: (2026)
by: Levinson, Matthew
Published: (2026)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Sanity Checks Revisited: An Exploration to Repair the Model Parameter Randomisation Test
by: Hedström, Anna, et al.
Published: (2024)
by: Hedström, Anna, et al.
Published: (2024)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
by: Roytburg, Dani, et al.
Published: (2026)
by: Roytburg, Dani, et al.
Published: (2026)
Sanity Checking Causal Representation Learning on a Simple Real-World System
by: Gamella, Juan L., et al.
Published: (2025)
by: Gamella, Juan L., et al.
Published: (2025)
Quasi-Random Physics-informed Neural Networks
by: Yu, Tianchi, et al.
Published: (2025)
by: Yu, Tianchi, et al.
Published: (2025)
Logit-KL Flow Matching: Non-Autoregressive Text Generation via Sampling-Hybrid Inference
by: Sevriugov, Egor, et al.
Published: (2024)
by: Sevriugov, Egor, et al.
Published: (2024)
Message-Passing GNNs Fail to Approximate Sparse Triangular Factorizations
by: Trifonov, Vladislav, et al.
Published: (2025)
by: Trifonov, Vladislav, et al.
Published: (2025)
The Rate-Distortion-Polysemanticity Tradeoff in SAEs
by: Mencattini, Tommaso, et al.
Published: (2026)
by: Mencattini, Tommaso, et al.
Published: (2026)
Tokenized SAEs: Disentangling SAE Reconstructions
by: Dooms, Thomas, et al.
Published: (2025)
by: Dooms, Thomas, et al.
Published: (2025)
Distribution-Aware Feature Selection for SAEs
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Sparse and Transferable Universal Singular Vectors Attack
by: Kuvshinova, Kseniia, et al.
Published: (2024)
by: Kuvshinova, Kseniia, et al.
Published: (2024)
Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control
by: Klenitskiy, Anton, et al.
Published: (2025)
by: Klenitskiy, Anton, et al.
Published: (2025)
Routing Absorption in Sparse Attention: Why Random Gates Are Hard to Beat
by: Aquino-Michaels, Keston
Published: (2026)
by: Aquino-Michaels, Keston
Published: (2026)
Analyzing (In)Abilities of SAEs via Formal Languages
by: Menon, Abhinav, et al.
Published: (2024)
by: Menon, Abhinav, et al.
Published: (2024)
Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities
by: Li, Pengyi, et al.
Published: (2026)
by: Li, Pengyi, et al.
Published: (2026)
Inverted Activations: Reducing Memory Footprint in Neural Network Training
by: Novikov, Georgii, et al.
Published: (2024)
by: Novikov, Georgii, et al.
Published: (2024)
Do Sparse Autoencoders Capture Concept Manifolds?
by: Bhalla, Usha, et al.
Published: (2026)
by: Bhalla, Usha, et al.
Published: (2026)
One Task Vector is not Enough: A Large-Scale Study for In-Context Learning
by: Tikhonov, Pavel, et al.
Published: (2025)
by: Tikhonov, Pavel, et al.
Published: (2025)
Do Sparse Autoencoders Identify Reasoning Features in Language Models?
by: Ma, George, et al.
Published: (2026)
by: Ma, George, et al.
Published: (2026)
Do Sparse Autoencoders Generalize? A Case Study of Answerability
by: Heindrich, Lovis, et al.
Published: (2025)
by: Heindrich, Lovis, et al.
Published: (2025)
Marchuk: Efficient Global Weather Forecasting from Mid-Range to Sub-Seasonal Scales via Flow Matching
by: Kuzhamuratov, Arsen, et al.
Published: (2026)
by: Kuzhamuratov, Arsen, et al.
Published: (2026)
Tensor-Train Point Cloud Compression and Efficient Approximate Nearest-Neighbor Search
by: Novikov, Georgii, et al.
Published: (2024)
by: Novikov, Georgii, et al.
Published: (2024)
Sparse Autoencoders Do Not Find Canonical Units of Analysis
by: Leask, Patrick, et al.
Published: (2025)
by: Leask, Patrick, et al.
Published: (2025)
Residual Stream Analysis with Multi-Layer SAEs
by: Lawson, Tim, et al.
Published: (2024)
by: Lawson, Tim, et al.
Published: (2024)
The benefits of query-based KGQA systems for complex and temporal questions in LLM era
by: Alekseev, Artem, et al.
Published: (2025)
by: Alekseev, Artem, et al.
Published: (2025)
Similar Items
-
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
by: Korznikov, Anton, et al.
Published: (2025) -
The Rogue Scalpel: Activation Steering Compromises LLM Safety
by: Korznikov, Anton, et al.
Published: (2025) -
I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
by: Galichin, Andrey, et al.
Published: (2025) -
GLiRA: Black-Box Membership Inference Attack via Knowledge Distillation
by: Galichin, Andrey V., et al.
Published: (2024) -
Spread them Apart: Towards Robust Watermarking of Generated Content
by: Pautov, Mikhail, et al.
Published: (2025)