Do Sparse Autoencoders Generalize? A Case Study of Answerability
Fuente:
arXiv
Saved in:
| Main Authors: | Heindrich, Lovis, Torr, Philip, Barez, Fazl, Thost, Veronika |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
by: Lan, Michael, et al.
Published: (2024)
by: Lan, Michael, et al.
Published: (2024)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)
by: Lan, Michael, et al.
Published: (2023)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
by: Marks, Luke, et al.
Published: (2024)
by: Marks, Luke, et al.
Published: (2024)
Understanding Addition in Transformers
by: Quirke, Philip, et al.
Published: (2023)
by: Quirke, Philip, et al.
Published: (2023)
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
by: Oldfield, James, et al.
Published: (2025)
by: Oldfield, James, et al.
Published: (2025)
Understanding Addition and Subtraction in Transformers
by: Quirke, Philip, et al.
Published: (2024)
by: Quirke, Philip, et al.
Published: (2024)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
by: Fu, Tingchen, et al.
Published: (2025)
by: Fu, Tingchen, et al.
Published: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Towards Interpreting Visual Information Processing in Vision-Language Models
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Interpreting Learned Feedback Patterns in Large Language Models
by: Marks, Luke, et al.
Published: (2023)
by: Marks, Luke, et al.
Published: (2023)
Self-Supervised Learning on Molecular Graphs: A Systematic Investigation of Masking Design
by: Yang, Jiannan, et al.
Published: (2025)
by: Yang, Jiannan, et al.
Published: (2025)
SAGE: Scalable Ground Truth Evaluations for Large Sparse Autoencoders
by: Venhoff, Constantin, et al.
Published: (2024)
by: Venhoff, Constantin, et al.
Published: (2024)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
by: Kim, Minseon, et al.
Published: (2025)
by: Kim, Minseon, et al.
Published: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
by: Schrodi, Simon, et al.
Published: (2025)
by: Schrodi, Simon, et al.
Published: (2025)
Visualizing Neural Network Imagination
by: Wichers, Nevan, et al.
Published: (2024)
by: Wichers, Nevan, et al.
Published: (2024)
Discovering the curriculum with AI: A proof-of-concept demonstration with an intelligent tutoring system for teaching project selection
by: Heindrich, Lovis, et al.
Published: (2024)
by: Heindrich, Lovis, et al.
Published: (2024)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Scaling sparse feature circuit finding for in-context learning
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
Do Sparse Autoencoders Capture Concept Manifolds?
by: Bhalla, Usha, et al.
Published: (2026)
by: Bhalla, Usha, et al.
Published: (2026)
Do Sparse Autoencoders Identify Reasoning Features in Language Models?
by: Ma, George, et al.
Published: (2026)
by: Ma, George, et al.
Published: (2026)
Sparse Autoencoders Do Not Find Canonical Units of Analysis
by: Leask, Patrick, et al.
Published: (2025)
by: Leask, Patrick, et al.
Published: (2025)
When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations
by: Petrov, Aleksandar, et al.
Published: (2023)
by: Petrov, Aleksandar, et al.
Published: (2023)
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
by: Korznikov, Anton, et al.
Published: (2026)
by: Korznikov, Anton, et al.
Published: (2026)
An intelligent tutor for planning in large partially observable environments
by: Heindrich, Lovis, et al.
Published: (2023)
by: Heindrich, Lovis, et al.
Published: (2023)
Ensembling Sparse Autoencoders
by: Gadgil, Soham, et al.
Published: (2025)
by: Gadgil, Soham, et al.
Published: (2025)
Scaling the Convex Barrier with Sparse Dual Algorithms
by: De Palma, Alessandro, et al.
Published: (2021)
by: De Palma, Alessandro, et al.
Published: (2021)
Analysis of Variational Sparse Autoencoders
by: Baker, Zachary, et al.
Published: (2025)
by: Baker, Zachary, et al.
Published: (2025)
Toward Identifiable Sparse Autoencoders
by: Nelson, Walter, et al.
Published: (2026)
by: Nelson, Walter, et al.
Published: (2026)
Sparse Autoencoders, Again?
by: Lu, Yin, et al.
Published: (2025)
by: Lu, Yin, et al.
Published: (2025)
Geometric Deviation as an Unsupervised Pre-Generation Reliability Signal: Probing LLM Representations for Answerability
by: Du, Yucheng
Published: (2026)
by: Du, Yucheng
Published: (2026)
Transcoders Beat Sparse Autoencoders for Interpretability
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Evaluating Sparse Autoencoders for Monosemantic Representation
by: Fereidouni, Moghis, et al.
Published: (2025)
by: Fereidouni, Moghis, et al.
Published: (2025)
Decomposing The Dark Matter of Sparse Autoencoders
by: Engels, Joshua, et al.
Published: (2024)
by: Engels, Joshua, et al.
Published: (2024)
Disentangling Dense Embeddings with Sparse Autoencoders
by: O'Neill, Charles, et al.
Published: (2024)
by: O'Neill, Charles, et al.
Published: (2024)
Are Sparse Autoencoder Benchmarks Reliable?
by: Chanin, David
Published: (2026)
by: Chanin, David
Published: (2026)
Low-Rank Adapting Models for Sparse Autoencoders
by: Chen, Matthew, et al.
Published: (2025)
by: Chen, Matthew, et al.
Published: (2025)
Attribution-Guided Distillation of Matryoshka Sparse Autoencoders
by: Martin-Linares, Cristina P., et al.
Published: (2025)
by: Martin-Linares, Cristina P., et al.
Published: (2025)
Interpretable Reward Model via Sparse Autoencoder
by: Zhang, Shuyi, et al.
Published: (2025)
by: Zhang, Shuyi, et al.
Published: (2025)
Similar Items
-
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
by: Lan, Michael, et al.
Published: (2024) -
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023) -
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
by: Marks, Luke, et al.
Published: (2024) -
Understanding Addition in Transformers
by: Quirke, Philip, et al.
Published: (2023) -
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
by: Oldfield, James, et al.
Published: (2025)