SAE-FD: Sparse Autoencoder Feature Distillation for Continual Learning of Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Mingxu, Li, Yuhan, Li, Lujundong, Shen, Dazhong, Xiong, Hui, Sun, Ying |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing
di: Zhang, Mingxu, et al.
Pubblicazione: (2026)
di: Zhang, Mingxu, et al.
Pubblicazione: (2026)
IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage
di: Li, Yuhan, et al.
Pubblicazione: (2026)
di: Li, Yuhan, et al.
Pubblicazione: (2026)
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
di: Cao, Tue M., et al.
Pubblicazione: (2026)
di: Cao, Tue M., et al.
Pubblicazione: (2026)
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
di: Korznikov, Anton, et al.
Pubblicazione: (2025)
di: Korznikov, Anton, et al.
Pubblicazione: (2025)
PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding
di: Koromilas, Panagiotis, et al.
Pubblicazione: (2026)
di: Koromilas, Panagiotis, et al.
Pubblicazione: (2026)
AlignSAE: Concept-Aligned Sparse Autoencoders
di: Yang, Minglai, et al.
Pubblicazione: (2025)
di: Yang, Minglai, et al.
Pubblicazione: (2025)
NGTM: Substructure-based Neural Graph Topic Model for Interpretable Graph Generation
di: Zhuang, Yuanxin, et al.
Pubblicazione: (2025)
di: Zhuang, Yuanxin, et al.
Pubblicazione: (2025)
SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection
di: Zhang, Huopu, et al.
Pubblicazione: (2025)
di: Zhang, Huopu, et al.
Pubblicazione: (2025)
Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE)
di: Brzozowski, Michał, et al.
Pubblicazione: (2026)
di: Brzozowski, Michał, et al.
Pubblicazione: (2026)
FD-LLM: Large Language Model for Fault Diagnosis of Machines
di: Qaid, Hamzah A. A. M., et al.
Pubblicazione: (2024)
di: Qaid, Hamzah A. A. M., et al.
Pubblicazione: (2024)
MolEditRL: Structure-Preserving Molecular Editing via Discrete Diffusion and Reinforcement Learning
di: Zhuang, Yuanxin, et al.
Pubblicazione: (2025)
di: Zhuang, Yuanxin, et al.
Pubblicazione: (2025)
Route Sparse Autoencoder to Interpret Large Language Models
di: Shi, Wei, et al.
Pubblicazione: (2025)
di: Shi, Wei, et al.
Pubblicazione: (2025)
ChemATP: A Training-Free Chemical Reasoning Framework for Large Language Models
di: Zhang, Mingxu, et al.
Pubblicazione: (2025)
di: Zhang, Mingxu, et al.
Pubblicazione: (2025)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
di: Cho, Seonglae, et al.
Pubblicazione: (2025)
di: Cho, Seonglae, et al.
Pubblicazione: (2025)
ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions
di: Poduval, Prathyush, et al.
Pubblicazione: (2026)
di: Poduval, Prathyush, et al.
Pubblicazione: (2026)
Do Sparse Autoencoders Identify Reasoning Features in Language Models?
di: Ma, George, et al.
Pubblicazione: (2026)
di: Ma, George, et al.
Pubblicazione: (2026)
Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models
di: Demircan, Can, et al.
Pubblicazione: (2024)
di: Demircan, Can, et al.
Pubblicazione: (2024)
Attribution-Guided Continual Learning for Large Language Models
di: Liu, Yazheng, et al.
Pubblicazione: (2026)
di: Liu, Yazheng, et al.
Pubblicazione: (2026)
SoftSAE: Dynamic Top-K Selection for Adaptive Sparse Autoencoders
di: Stępień, Jakub, et al.
Pubblicazione: (2026)
di: Stępień, Jakub, et al.
Pubblicazione: (2026)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
di: Pach, Mateusz, et al.
Pubblicazione: (2025)
di: Pach, Mateusz, et al.
Pubblicazione: (2025)
Attribution-Guided Distillation of Matryoshka Sparse Autoencoders
di: Martin-Linares, Cristina P., et al.
Pubblicazione: (2025)
di: Martin-Linares, Cristina P., et al.
Pubblicazione: (2025)
The Geometry of Concepts: Sparse Autoencoder Feature Structure
di: Li, Yuxiao, et al.
Pubblicazione: (2024)
di: Li, Yuxiao, et al.
Pubblicazione: (2024)
Distillation-Guided Structural Transfer for Continual Learning Beyond Sparse Distributed Memory
di: Xue, Huiyan, et al.
Pubblicazione: (2025)
di: Xue, Huiyan, et al.
Pubblicazione: (2025)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
di: Lan, Michael, et al.
Pubblicazione: (2024)
di: Lan, Michael, et al.
Pubblicazione: (2024)
AtomDisc: An Atom-level Tokenizer that Boosts Molecular LLMs and Reveals Structure--Property Associations
di: Zhang, Mingxu, et al.
Pubblicazione: (2025)
di: Zhang, Mingxu, et al.
Pubblicazione: (2025)
Sparse Autoencoders Trained on the Same Data Learn Different Features
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
di: Paulo, Gonçalo, et al.
Pubblicazione: (2025)
MO-SAE:Multi-Objective Stacked Autoencoders Optimization for Edge Anomaly Detection
di: Zhang, Lizhao, et al.
Pubblicazione: (2026)
di: Zhang, Lizhao, et al.
Pubblicazione: (2026)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
di: Bussmann, Bart, et al.
Pubblicazione: (2025)
di: Bussmann, Bart, et al.
Pubblicazione: (2025)
Interpretable Reward Model via Sparse Autoencoder
di: Zhang, Shuyi, et al.
Pubblicazione: (2025)
di: Zhang, Shuyi, et al.
Pubblicazione: (2025)
Steering Language Model Refusal with Sparse Autoencoders
di: O'Brien, Kyle, et al.
Pubblicazione: (2024)
di: O'Brien, Kyle, et al.
Pubblicazione: (2024)
Analysis of Variational Sparse Autoencoders
di: Baker, Zachary, et al.
Pubblicazione: (2025)
di: Baker, Zachary, et al.
Pubblicazione: (2025)
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
di: Yamashita, Tomoya, et al.
Pubblicazione: (2025)
di: Yamashita, Tomoya, et al.
Pubblicazione: (2025)
Enhance the Safety in Reinforcement Learning by ADRC Lagrangian Methods
di: Zhang, Mingxu, et al.
Pubblicazione: (2026)
di: Zhang, Mingxu, et al.
Pubblicazione: (2026)
Adaptive Regularization for Large-Scale Sparse Feature Embedding Models
di: Li, Mang, et al.
Pubblicazione: (2025)
di: Li, Mang, et al.
Pubblicazione: (2025)
Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
di: Paek, Nathan, et al.
Pubblicazione: (2025)
di: Paek, Nathan, et al.
Pubblicazione: (2025)
GeoSAE: Geometric Prior-Guided Layer-Wise Sparse Autoencoder Annotation of Brain MRI Foundation Models
di: Nerrise, Favour, et al.
Pubblicazione: (2026)
di: Nerrise, Favour, et al.
Pubblicazione: (2026)
Training Superior Sparse Autoencoders for Instruct Models
di: Li, Jiaming, et al.
Pubblicazione: (2025)
di: Li, Jiaming, et al.
Pubblicazione: (2025)
Dense SAE Latents Are Features, Not Bugs
di: Sun, Xiaoqing, et al.
Pubblicazione: (2025)
di: Sun, Xiaoqing, et al.
Pubblicazione: (2025)
Improving Sparse Autoencoder with Dynamic Attention
di: Wang, Dongsheng, et al.
Pubblicazione: (2026)
di: Wang, Dongsheng, et al.
Pubblicazione: (2026)
Evolution of SAE Features Across Layers in LLMs
di: Balcells, Daniel, et al.
Pubblicazione: (2024)
di: Balcells, Daniel, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing
di: Zhang, Mingxu, et al.
Pubblicazione: (2026) -
IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage
di: Li, Yuhan, et al.
Pubblicazione: (2026) -
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
di: Cao, Tue M., et al.
Pubblicazione: (2026) -
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
di: Korznikov, Anton, et al.
Pubblicazione: (2025) -
PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding
di: Koromilas, Panagiotis, et al.
Pubblicazione: (2026)