Understanding sparse autoencoder scaling in the presence of feature manifolds
Fuente:
arXiv
Saved in:
| Main Authors: | Michaud, Eric J., Gorton, Liv, McGrath, Tom |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Group Crosscoders for Mechanistic Analysis of Symmetry
by: Gorton, Liv
Published: (2024)
by: Gorton, Liv
Published: (2024)
The Missing Curve Detectors of InceptionV1: Applying Sparse Autoencoders to InceptionV1 Early Vision
by: Gorton, Liv
Published: (2024)
by: Gorton, Liv
Published: (2024)
Adversarial Examples Are Not Bugs, They Are Superposition
by: Gorton, Liv, et al.
Published: (2025)
by: Gorton, Liv, et al.
Published: (2025)
Scaling and evaluating sparse autoencoders
by: Gao, Leo, et al.
Published: (2024)
by: Gao, Leo, et al.
Published: (2024)
Pairwise matrices for sparse autoencoders: single-feature inspection mislabels causal axes
by: Riegler, Michael A., et al.
Published: (2026)
by: Riegler, Michael A., et al.
Published: (2026)
Nuisance Function Tuning and Sample Splitting for Optimally Estimating a Doubly Robust Functional
by: McGrath, Sean, et al.
Published: (2022)
by: McGrath, Sean, et al.
Published: (2022)
Shifting the Gradient: Understanding How Defensive Training Methods Protect Language Model Integrity
by: Grant, Satchel, et al.
Published: (2026)
by: Grant, Satchel, et al.
Published: (2026)
Decomposing multimodal embedding spaces with group-sparse autoencoders
by: Kaushik, Chiraag, et al.
Published: (2026)
by: Kaushik, Chiraag, et al.
Published: (2026)
Bilinear autoencoders find interpretable manifolds
by: Dooms, Thomas, et al.
Published: (2026)
by: Dooms, Thomas, et al.
Published: (2026)
Investigating task-specific prompts and sparse autoencoders for activation monitoring
by: Tillman, Henk, et al.
Published: (2025)
by: Tillman, Henk, et al.
Published: (2025)
Applying sparse autoencoders to unlearn knowledge in language models
by: Farrell, Eoin, et al.
Published: (2024)
by: Farrell, Eoin, et al.
Published: (2024)
Steering CLIP's vision transformer with sparse autoencoders
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
Insights into a radiology-specialised multimodal large language model with sparse autoencoders
by: Bouzid, Kenza, et al.
Published: (2025)
by: Bouzid, Kenza, et al.
Published: (2025)
Can sparse autoencoders make sense of gene expression latent variable models?
by: Schuster, Viktoria
Published: (2024)
by: Schuster, Viktoria
Published: (2024)
Measuring the Representational Alignment of Neural Systems in Superposition
by: Liu, Sunny, et al.
Published: (2026)
by: Liu, Sunny, et al.
Published: (2026)
Can sparse autoencoders be used to decompose and interpret steering vectors?
by: Mayne, Harry, et al.
Published: (2024)
by: Mayne, Harry, et al.
Published: (2024)
Ransomware detection using stacked autoencoder for feature selection
by: Nkongolo, Mike, et al.
Published: (2024)
by: Nkongolo, Mike, et al.
Published: (2024)
From Frege to chatGPT: Compositionality in language, cognition, and deep neural networks
by: Russin, Jacob, et al.
Published: (2024)
by: Russin, Jacob, et al.
Published: (2024)
MATRIX: A Multimodal Benchmark and Post-Training Framework for Materials Science
by: McGrath, Delia, et al.
Published: (2026)
by: McGrath, Delia, et al.
Published: (2026)
Transformer autoencoder with local attention for sparse and irregular time series with application on risk estimation
by: Rodis, Panteleimon
Published: (2026)
by: Rodis, Panteleimon
Published: (2026)
Slim multi-scale convolutional autoencoder-based reduced-order models for interpretable features of a complex dynamical system
by: Teutsch, Philipp, et al.
Published: (2025)
by: Teutsch, Philipp, et al.
Published: (2025)
Optimal Nuisance Function Tuning for Estimating a Doubly Robust Functional under Proportional Asymptotics
by: McGrath, Sean, et al.
Published: (2025)
by: McGrath, Sean, et al.
Published: (2025)
On the creation of narrow AI: hierarchy and nonlocality of neural network skills
by: Michaud, Eric J., et al.
Published: (2025)
by: Michaud, Eric J., et al.
Published: (2025)
LEARNER: A Transfer Learning Method for Low-Rank Matrix Estimation
by: McGrath, Sean, et al.
Published: (2024)
by: McGrath, Sean, et al.
Published: (2024)
DeepAtlas: a tool for effective manifold learning
by: Hughes, Serena, et al.
Published: (2025)
by: Hughes, Serena, et al.
Published: (2025)
Understanding the role of autoencoders for stiff dynamical systems using information theory
by: Vijayarangan, Vijayamanikandan, et al.
Published: (2025)
by: Vijayarangan, Vijayamanikandan, et al.
Published: (2025)
Why should autoencoders work?
by: Kvalheim, Matthew D., et al.
Published: (2023)
by: Kvalheim, Matthew D., et al.
Published: (2023)
Lessons Jesuit Business Programs Can Learn from Chinese MBA Programs
by: Mary Ann McGrath
Published: (2016)
by: Mary Ann McGrath
Published: (2016)
Scaling sparse feature circuit finding for in-context learning
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability
by: Prasad, Aaditya Vikram, et al.
Published: (2026)
by: Prasad, Aaditya Vikram, et al.
Published: (2026)
Matching aggregate posteriors in the variational autoencoder
by: Saha, Surojit, et al.
Published: (2023)
by: Saha, Surojit, et al.
Published: (2023)
Not All Language Model Features Are One-Dimensionally Linear
by: Engels, Joshua, et al.
Published: (2024)
by: Engels, Joshua, et al.
Published: (2024)
Learning-based estimation of cattle weight gain and its influencing factors
by: Hossain, Muhammad Riaz Hasib, et al.
Published: (2025)
by: Hossain, Muhammad Riaz Hasib, et al.
Published: (2025)
Mob-based cattle weight gain forecasting using ML models
by: Hossain, Muhammad Riaz Hasib, et al.
Published: (2025)
by: Hossain, Muhammad Riaz Hasib, et al.
Published: (2025)
Global atmospheric data assimilation with multi-modal masked autoencoders
by: Vandal, Thomas J., et al.
Published: (2024)
by: Vandal, Thomas J., et al.
Published: (2024)
Predicting large scale cosmological structure evolution with generative adversarial network-based autoencoders
by: Ullmo, Marion, et al.
Published: (2024)
by: Ullmo, Marion, et al.
Published: (2024)
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
by: Bigelow, Eric, et al.
Published: (2026)
by: Bigelow, Eric, et al.
Published: (2026)
Complex variational autoencoders admit Kähler structure
by: Gracyk, Andrew
Published: (2025)
by: Gracyk, Andrew
Published: (2025)
Accelerating spherical K-means clustering for large-scale sparse document data
by: Aoyama, Kazuo, et al.
Published: (2024)
by: Aoyama, Kazuo, et al.
Published: (2024)
Learning Environment for the Air Domain (LEAD)
by: Strand, Andreas, et al.
Published: (2023)
by: Strand, Andreas, et al.
Published: (2023)
Similar Items
-
Group Crosscoders for Mechanistic Analysis of Symmetry
by: Gorton, Liv
Published: (2024) -
The Missing Curve Detectors of InceptionV1: Applying Sparse Autoencoders to InceptionV1 Early Vision
by: Gorton, Liv
Published: (2024) -
Adversarial Examples Are Not Bugs, They Are Superposition
by: Gorton, Liv, et al.
Published: (2025) -
Scaling and evaluating sparse autoencoders
by: Gao, Leo, et al.
Published: (2024) -
Pairwise matrices for sparse autoencoders: single-feature inspection mislabels causal axes
by: Riegler, Michael A., et al.
Published: (2026)