On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted Remedy
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Jingyi, Zhang, Qi, Wang, Yifei, Wang, Yisen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Inclusive Theoretical Framework of Robust Supervised Contrastive Loss against Label Noise
by: Cui, Jingyi, et al.
Published: (2025)
by: Cui, Jingyi, et al.
Published: (2025)
Difficult Examples Hurt Unsupervised Contrastive Learning: A Theoretical Perspective
by: Zhang, Yi-Ge, et al.
Published: (2025)
by: Zhang, Yi-Ge, et al.
Published: (2025)
Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
An Augmentation-Aware Theory for Self-Supervised Contrastive Learning
by: Cui, Jingyi, et al.
Published: (2025)
by: Cui, Jingyi, et al.
Published: (2025)
An Augmentation Overlap Theory of Contrastive Learning
by: Zhang, Qi, et al.
Published: (2025)
by: Zhang, Qi, et al.
Published: (2025)
A Theoretical Understanding of Self-Correction through In-context Alignment
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Laplacian Canonization: A Minimalist Approach to Sign and Basis Invariant Spectral Embedding
by: Ma, Jiangyan, et al.
Published: (2023)
by: Ma, Jiangyan, et al.
Published: (2023)
Non-negative Contrastive Learning
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Projection Head is Secretly an Information Bottleneck
by: Ouyang, Zhuo, et al.
Published: (2025)
by: Ouyang, Zhuo, et al.
Published: (2025)
Leveraging Flatness to Improve Information-Theoretic Generalization Bounds for SGD
by: Peng, Ze, et al.
Published: (2026)
by: Peng, Ze, et al.
Published: (2026)
Do Generated Data Always Help Contrastive Learning?
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations
by: Yao, Yifei, et al.
Published: (2025)
by: Yao, Yifei, et al.
Published: (2025)
On the Role of Discrete Tokenization in Visual Representation Learning
by: Du, Tianqi, et al.
Published: (2024)
by: Du, Tianqi, et al.
Published: (2024)
DFedReweighting: A Unified Framework for Objective-Oriented Reweighting in Decentralized Federated Learning
by: Zhang, Kaichuang, et al.
Published: (2025)
by: Zhang, Kaichuang, et al.
Published: (2025)
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
Dissecting the Failure of Invariant Learning on Graphs
by: Wang, Qixun, et al.
Published: (2024)
by: Wang, Qixun, et al.
Published: (2024)
Can In-context Learning Really Generalize to Out-of-distribution Tasks?
by: Wang, Qixun, et al.
Published: (2024)
by: Wang, Qixun, et al.
Published: (2024)
A Canonicalization Perspective on Invariant and Equivariant Learning
by: Ma, George, et al.
Published: (2024)
by: Ma, George, et al.
Published: (2024)
Long-Short Alignment for Effective Long-Context Modeling in LLMs
by: Du, Tianqi, et al.
Published: (2025)
by: Du, Tianqi, et al.
Published: (2025)
An Iteratively Reweighted Method for Sparse Optimization on Nonconvex $\ell_{p}$ Ball
by: Wang, Hao, et al.
Published: (2021)
by: Wang, Hao, et al.
Published: (2021)
Improving Sparse Autoencoder with Dynamic Attention
by: Wang, Dongsheng, et al.
Published: (2026)
by: Wang, Dongsheng, et al.
Published: (2026)
Interpretable Reward Model via Sparse Autoencoder
by: Zhang, Shuyi, et al.
Published: (2025)
by: Zhang, Shuyi, et al.
Published: (2025)
How to Craft Backdoors with Unlabeled Data Alone?
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Adversarial Examples Are Not Real Features
by: Li, Ang, et al.
Published: (2023)
by: Li, Ang, et al.
Published: (2023)
Revisiting Meta-Learning with Noisy Labels: Reweighting Dynamics and Theoretical Guarantees
by: Zhang, Yiming, et al.
Published: (2025)
by: Zhang, Yiming, et al.
Published: (2025)
Iterative Reweighted Framework Based Algorithms for Sparse Linear Regression with Generalized Elastic Net Penalty
by: Ding, Yanyun, et al.
Published: (2024)
by: Ding, Yanyun, et al.
Published: (2024)
LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
by: Guo, Xiaojun, et al.
Published: (2025)
by: Guo, Xiaojun, et al.
Published: (2025)
Ensembling Sparse Autoencoders
by: Gadgil, Soham, et al.
Published: (2025)
by: Gadgil, Soham, et al.
Published: (2025)
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
by: Cao, Qi, et al.
Published: (2025)
by: Cao, Qi, et al.
Published: (2025)
FedOAED: Federated On-Device Autoencoder Denoiser for Heterogeneous Data under Limited Client Availability
by: Howlader, S M Ruhul Kabir, et al.
Published: (2025)
by: Howlader, S M Ruhul Kabir, et al.
Published: (2025)
Analysis of Variational Sparse Autoencoders
by: Baker, Zachary, et al.
Published: (2025)
by: Baker, Zachary, et al.
Published: (2025)
Toward Identifiable Sparse Autoencoders
by: Nelson, Walter, et al.
Published: (2026)
by: Nelson, Walter, et al.
Published: (2026)
Physics-guided Active Sample Reweighting for Urban Flow Prediction
by: Jiang, Wei, et al.
Published: (2024)
by: Jiang, Wei, et al.
Published: (2024)
Route Sparse Autoencoder to Interpret Large Language Models
by: Shi, Wei, et al.
Published: (2025)
by: Shi, Wei, et al.
Published: (2025)
Topology-Aware Dynamic Reweighting for Distribution Shifts on Graph
by: Zheng, Weihuang, et al.
Published: (2024)
by: Zheng, Weihuang, et al.
Published: (2024)
DAJ: Data-Reweighted LLM Judge for Test-Time Scaling in Code Generation
by: Qin, Peijia, et al.
Published: (2026)
by: Qin, Peijia, et al.
Published: (2026)
Sparse Autoencoders, Again?
by: Lu, Yin, et al.
Published: (2025)
by: Lu, Yin, et al.
Published: (2025)
Theoretical Convergence Guarantees for Variational Autoencoders
by: Surendran, Sobihan, et al.
Published: (2024)
by: Surendran, Sobihan, et al.
Published: (2024)
Similar Items
-
An Inclusive Theoretical Framework of Robust Supervised Contrastive Loss against Label Noise
by: Cui, Jingyi, et al.
Published: (2025) -
Difficult Examples Hurt Unsupervised Contrastive Learning: A Theoretical Perspective
by: Zhang, Yi-Ge, et al.
Published: (2025) -
Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining
by: Zhang, Qi, et al.
Published: (2024) -
Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
by: Zhang, Qi, et al.
Published: (2024) -
An Augmentation-Aware Theory for Self-Supervised Contrastive Learning
by: Cui, Jingyi, et al.
Published: (2025)