Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Siyu, Sheen, Heejune, Xiong, Xuyuan, Wang, Tianhao, Yang, Zhuoran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2026)
by: Cho, Seonglae, et al.
Published: (2026)
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2025)
by: Cho, Seonglae, et al.
Published: (2025)
The Information-Theoretic Imperative: Compression and the Epistemic Foundations of Intelligence
by: Dittrich, Christian, et al.
Published: (2025)
by: Dittrich, Christian, et al.
Published: (2025)
Learning to Construct Knowledge through Sparse Reference Selection with Reinforcement Learning
by: Yin, Shao-An
Published: (2025)
by: Yin, Shao-An
Published: (2025)
Task Adaptation from Skills: Information Geometry, Disentanglement, and New Objectives for Unsupervised Reinforcement Learning
by: Yang, Yucheng, et al.
Published: (2025)
by: Yang, Yucheng, et al.
Published: (2025)
Inference Time Causal Probing in LLMs
by: Khorasani, Sadegh, et al.
Published: (2026)
by: Khorasani, Sadegh, et al.
Published: (2026)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
by: Borobia, Hector, et al.
Published: (2026)
by: Borobia, Hector, et al.
Published: (2026)
Loss-Complexity Landscape and Model Structure Functions
by: Kolpakov, Alexander
Published: (2025)
by: Kolpakov, Alexander
Published: (2025)
ATANT v1.1: Positioning Continuity Evaluation Against Memory, Long-Context, and Agentic-Memory Benchmarks
by: Tanguturi, Samuel Sameer
Published: (2026)
by: Tanguturi, Samuel Sameer
Published: (2026)
ATANT: An Evaluation Framework for AI Continuity
by: Tanguturi, Samuel Sameer
Published: (2026)
by: Tanguturi, Samuel Sameer
Published: (2026)
Understanding Variational Autoencoders with Intrinsic Dimension and Information Imbalance
by: Camboulin, Charles, et al.
Published: (2024)
by: Camboulin, Charles, et al.
Published: (2024)
Tensor Generalized Approximate Message Passing
by: Li, Yinchuan, et al.
Published: (2025)
by: Li, Yinchuan, et al.
Published: (2025)
Beyond Mimicry: Preference Coherence in LLMs
by: Mikaelson, Luhan, et al.
Published: (2025)
by: Mikaelson, Luhan, et al.
Published: (2025)
Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery
by: Yuan, Xinzhe, et al.
Published: (2026)
by: Yuan, Xinzhe, et al.
Published: (2026)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
AI Benchmark Democratization and Carpentry
by: von Laszewski, Gregor, et al.
Published: (2025)
by: von Laszewski, Gregor, et al.
Published: (2025)
Accelerating Complex Disease Treatment through Network Medicine and GenAI: A Case Study on Drug Repurposing for Breast Cancer
by: Hamed, Ahmed Abdeen, et al.
Published: (2024)
by: Hamed, Ahmed Abdeen, et al.
Published: (2024)
Constrained Auto-Bidding via Generative Response Modeling
by: Yang, Eunseok, et al.
Published: (2026)
by: Yang, Eunseok, et al.
Published: (2026)
ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing
by: Kadu, Ankush, et al.
Published: (2025)
by: Kadu, Ankush, et al.
Published: (2025)
CogRec: A Cognitive Recommender Agent Fusing Large Language Models and Soar for Explainable Recommendation
by: Hu, Jiaxin, et al.
Published: (2025)
by: Hu, Jiaxin, et al.
Published: (2025)
Teacher-Student Guided Inverse Modeling for Steel Final Hardness Estimation
by: Alsheikh, Ahmad, et al.
Published: (2025)
by: Alsheikh, Ahmad, et al.
Published: (2025)
Comprehensive Metapath-based Heterogeneous Graph Transformer for Gene-Disease Association Prediction
by: Cui, Wentao, et al.
Published: (2025)
by: Cui, Wentao, et al.
Published: (2025)
LLMs in the Loop: Leveraging Large Language Model Annotations for Active Learning in Low-Resource Languages
by: Kholodna, Nataliia, et al.
Published: (2024)
by: Kholodna, Nataliia, et al.
Published: (2024)
Let Your Graph Do the Talking: Encoding Structured Data for LLMs
by: Perozzi, Bryan, et al.
Published: (2024)
by: Perozzi, Bryan, et al.
Published: (2024)
Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Linear-Readout Floors and Threshold Recovery in Computation in Superposition
by: Borobia, Hector, et al.
Published: (2026)
by: Borobia, Hector, et al.
Published: (2026)
Introducing New Node Prediction in Graph Mining: Predicting All Links from Isolated Nodes with Graph Neural Networks
by: Zanardini, Damiano, et al.
Published: (2024)
by: Zanardini, Damiano, et al.
Published: (2024)
CogCanvas: Verbatim-Grounded Artifact Extraction for Long LLM Conversations
by: An, Tao
Published: (2025)
by: An, Tao
Published: (2025)
Graph Your Way to Inspiration: Integrating Co-Author Graphs with Retrieval-Augmented Generation for Large Language Model Based Scientific Idea Generation
by: Xie, Pengzhen, et al.
Published: (2025)
by: Xie, Pengzhen, et al.
Published: (2025)
Perceptron Collaborative Filtering
by: Chakraborty, Arya
Published: (2024)
by: Chakraborty, Arya
Published: (2024)
Sketch Decompositions for Classical Planning via Deep Reinforcement Learning
by: Aichmüller, Michael, et al.
Published: (2024)
by: Aichmüller, Michael, et al.
Published: (2024)
Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers
by: Avinash, Mynampati Sri Ranganadha
Published: (2026)
by: Avinash, Mynampati Sri Ranganadha
Published: (2026)
Robust and Diverse Multi-Agent Learning via Rational Policy Gradient
by: Lauffer, Niklas, et al.
Published: (2025)
by: Lauffer, Niklas, et al.
Published: (2025)
Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution
by: Liu, Xiaoou, et al.
Published: (2026)
by: Liu, Xiaoou, et al.
Published: (2026)
X-SHIELD: Regularization for eXplainable Artificial Intelligence
by: Sevillano-García, Iván, et al.
Published: (2024)
by: Sevillano-García, Iván, et al.
Published: (2024)
Transformer Mechanisms Mimic Frontostriatal Gating Operations When Trained on Human Working Memory Tasks
by: Traylor, Aaron, et al.
Published: (2024)
by: Traylor, Aaron, et al.
Published: (2024)
Scalable and Robust LLM Unlearning by Correcting Responses with Retrieved Exclusions
by: Kim, Junbeom, et al.
Published: (2025)
by: Kim, Junbeom, et al.
Published: (2025)
Collaborative Multi-Agent Scripts Generation for Enhancing Imperfect-Information Reasoning in Murder Mystery Games
by: Zhong, Keyang, et al.
Published: (2026)
by: Zhong, Keyang, et al.
Published: (2026)
Similar Items
-
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
by: Chen, Siyu, et al.
Published: (2024) -
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2026) -
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
by: Chen, Siyu, et al.
Published: (2024) -
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2025) -
The Information-Theoretic Imperative: Compression and the Epistemic Foundations of Intelligence
by: Dittrich, Christian, et al.
Published: (2025)