Sparse Autoencoders, Again?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Yin, Zhu, Xuening, He, Tong, Wipf, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Are Sparse Autoencoder Benchmarks Reliable?
von: Chanin, David
Veröffentlicht: (2026)
von: Chanin, David
Veröffentlicht: (2026)
Data Whitening Improves Sparse Autoencoder Learning
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025)
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025)
BatchTopK Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
Improving Sparse Autoencoder with Dynamic Attention
von: Wang, Dongsheng, et al.
Veröffentlicht: (2026)
von: Wang, Dongsheng, et al.
Veröffentlicht: (2026)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
Transformers from Diffusion: A Unified Framework for Neural Message Passing
von: Wu, Qitian, et al.
Veröffentlicht: (2024)
von: Wu, Qitian, et al.
Veröffentlicht: (2024)
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
von: Chanin, David, et al.
Veröffentlicht: (2026)
von: Chanin, David, et al.
Veröffentlicht: (2026)
Never Reset Again: A Mathematical Framework for Continual Inference in Recurrent Neural Networks
von: Yin, Bojian, et al.
Veröffentlicht: (2024)
von: Yin, Bojian, et al.
Veröffentlicht: (2024)
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
von: Ayonrinde, Kola
Veröffentlicht: (2024)
von: Ayonrinde, Kola
Veröffentlicht: (2024)
Towards Interpretable Protein Structure Prediction with Sparse Autoencoders
von: Parsan, Nithin, et al.
Veröffentlicht: (2025)
von: Parsan, Nithin, et al.
Veröffentlicht: (2025)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
von: Zhu, Xudong, et al.
Veröffentlicht: (2025)
von: Zhu, Xudong, et al.
Veröffentlicht: (2025)
Empirical Evaluation of Progressive Coding for Sparse Autoencoders
von: Peter, Hans, et al.
Veröffentlicht: (2025)
von: Peter, Hans, et al.
Veröffentlicht: (2025)
On the transferability of Sparse Autoencoders for interpreting compressed models
von: Gupte, Suchit, et al.
Veröffentlicht: (2025)
von: Gupte, Suchit, et al.
Veröffentlicht: (2025)
Do Sparse Autoencoders Capture Concept Manifolds?
von: Bhalla, Usha, et al.
Veröffentlicht: (2026)
von: Bhalla, Usha, et al.
Veröffentlicht: (2026)
Improving Dictionary Learning with Gated Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025)
von: Chanin, David, et al.
Veröffentlicht: (2025)
Born Again Neural Networks
von: Furlanello, Tommaso, et al.
Veröffentlicht: (2018)
von: Furlanello, Tommaso, et al.
Veröffentlicht: (2018)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
Sparse Autoencoders Do Not Find Canonical Units of Analysis
von: Leask, Patrick, et al.
Veröffentlicht: (2025)
von: Leask, Patrick, et al.
Veröffentlicht: (2025)
Evaluating and Designing Sparse Autoencoders by Approximating Quasi-Orthogonality
von: Lee, Sewoong, et al.
Veröffentlicht: (2025)
von: Lee, Sewoong, et al.
Veröffentlicht: (2025)
Graph-Regularized Sparse Autoencoders for LLM Safety Steering
von: Yeon, Jehyeok, et al.
Veröffentlicht: (2025)
von: Yeon, Jehyeok, et al.
Veröffentlicht: (2025)
Improving Robustness In Sparse Autoencoders via Masked Regularization
von: Narayanaswamy, Vivek, et al.
Veröffentlicht: (2026)
von: Narayanaswamy, Vivek, et al.
Veröffentlicht: (2026)
The Geometry of Concepts: Sparse Autoencoder Feature Structure
von: Li, Yuxiao, et al.
Veröffentlicht: (2024)
von: Li, Yuxiao, et al.
Veröffentlicht: (2024)
Handling Distribution Shifts on Graphs: An Invariance Perspective
von: Wu, Qitian, et al.
Veröffentlicht: (2022)
von: Wu, Qitian, et al.
Veröffentlicht: (2022)
Beyond Redundancy: Diverse and Specialized Multi-Expert Sparse Autoencoder
von: Xu, Zhen, et al.
Veröffentlicht: (2025)
von: Xu, Zhen, et al.
Veröffentlicht: (2025)
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
von: Hindupur, Sai Sumedh R., et al.
Veröffentlicht: (2025)
von: Hindupur, Sai Sumedh R., et al.
Veröffentlicht: (2025)
Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
von: Khoriaty, Matthew, et al.
Veröffentlicht: (2025)
von: Khoriaty, Matthew, et al.
Veröffentlicht: (2025)
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2025)
von: Chanin, David, et al.
Veröffentlicht: (2025)
Sparse Autoencoder Features for Classifications and Transferability
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
Learning Retrieval Models with Sparse Autoencoders
von: Formal, Thibault, et al.
Veröffentlicht: (2026)
von: Formal, Thibault, et al.
Veröffentlicht: (2026)
Finding Belief Geometries with Sparse Autoencoders
von: Levinson, Matthew
Veröffentlicht: (2026)
von: Levinson, Matthew
Veröffentlicht: (2026)
Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability
von: Bereska, Leonard, et al.
Veröffentlicht: (2025)
von: Bereska, Leonard, et al.
Veröffentlicht: (2025)
Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
von: Wang, Anyi, et al.
Veröffentlicht: (2025)
von: Wang, Anyi, et al.
Veröffentlicht: (2025)
SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions
von: Poduval, Prathyush, et al.
Veröffentlicht: (2026)
von: Poduval, Prathyush, et al.
Veröffentlicht: (2026)
Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
von: Cao, Tue M., et al.
Veröffentlicht: (2026)
A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
von: Liu, Licheng, et al.
Veröffentlicht: (2025)
von: Liu, Licheng, et al.
Veröffentlicht: (2025)
SparseDM: Toward Sparse Efficient Diffusion Models
von: Wang, Kafeng, et al.
Veröffentlicht: (2024)
von: Wang, Kafeng, et al.
Veröffentlicht: (2024)
Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
von: Zhao, Daniel, et al.
Veröffentlicht: (2025)
von: Zhao, Daniel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Are Sparse Autoencoder Benchmarks Reliable?
von: Chanin, David
Veröffentlicht: (2026) -
Data Whitening Improves Sparse Autoencoder Learning
von: Saraswatula, Ashwin, et al.
Veröffentlicht: (2025) -
BatchTopK Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2024) -
Improving Sparse Autoencoder with Dynamic Attention
von: Wang, Dongsheng, et al.
Veröffentlicht: (2026) -
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)