Salvato in:
| Autori principali: | Wang, Xu, Li, Zihao, Wang, Benyou, Hu, Yan, Zou, Difan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2505.24428 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2025)
di: Wang, Xu, et al.
Pubblicazione: (2025)
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
di: Wang, Xu, et al.
Pubblicazione: (2025)
di: Wang, Xu, et al.
Pubblicazione: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2026)
di: Wang, Xu, et al.
Pubblicazione: (2026)
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
di: Yamashita, Tomoya, et al.
Pubblicazione: (2025)
di: Yamashita, Tomoya, et al.
Pubblicazione: (2025)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
di: Jing, Yi, et al.
Pubblicazione: (2026)
di: Jing, Yi, et al.
Pubblicazione: (2026)
A Human-Like Reasoning Framework for Multi-Phases Planning Task with Large Language Models
di: Xie, Chengxing, et al.
Pubblicazione: (2024)
di: Xie, Chengxing, et al.
Pubblicazione: (2024)
Dissecting Fine-Tuning Unlearning in Large Language Models
di: Hong, Yihuai, et al.
Pubblicazione: (2024)
di: Hong, Yihuai, et al.
Pubblicazione: (2024)
Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
di: Yan, Xinyuan, et al.
Pubblicazione: (2025)
di: Yan, Xinyuan, et al.
Pubblicazione: (2025)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
Training Superior Sparse Autoencoders for Instruct Models
di: Li, Jiaming, et al.
Pubblicazione: (2025)
di: Li, Jiaming, et al.
Pubblicazione: (2025)
SOMP: Scalable Gradient Inversion for Large Language Models via Subspace-Guided Orthogonal Matching Pursuit
di: Li, Yibo, et al.
Pubblicazione: (2026)
di: Li, Yibo, et al.
Pubblicazione: (2026)
PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding
di: Koromilas, Panagiotis, et al.
Pubblicazione: (2026)
di: Koromilas, Panagiotis, et al.
Pubblicazione: (2026)
ProxyAttn: Guided Sparse Attention via Representative Heads
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
Rotation Control Unlearning: Quantifying and Controlling Continuous Unlearning for LLM with The Cognitive Rotation Space
di: Zhang, Xiang, et al.
Pubblicazione: (2025)
di: Zhang, Xiang, et al.
Pubblicazione: (2025)
Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
di: Zheng, Carolina, et al.
Pubblicazione: (2025)
di: Zheng, Carolina, et al.
Pubblicazione: (2025)
An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models
di: Hu, Yunzhe, et al.
Pubblicazione: (2024)
di: Hu, Yunzhe, et al.
Pubblicazione: (2024)
Boosting Protein Language Models with Negative Sample Mining
di: Xu, Yaoyao, et al.
Pubblicazione: (2024)
di: Xu, Yaoyao, et al.
Pubblicazione: (2024)
Reviving Your MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing
di: Kassem, Aly M., et al.
Pubblicazione: (2025)
di: Kassem, Aly M., et al.
Pubblicazione: (2025)
AlignSAE: Concept-Aligned Sparse Autoencoders
di: Yang, Minglai, et al.
Pubblicazione: (2025)
di: Yang, Minglai, et al.
Pubblicazione: (2025)
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
di: O'Neill, Charles, et al.
Pubblicazione: (2024)
di: O'Neill, Charles, et al.
Pubblicazione: (2024)
LLM Unlearning with LLM Beliefs
di: Li, Kemou, et al.
Pubblicazione: (2025)
di: Li, Kemou, et al.
Pubblicazione: (2025)
PrivacyScalpel: Enhancing LLM Privacy via Interpretable Feature Intervention with Sparse Autoencoders
di: Frikha, Ahmed, et al.
Pubblicazione: (2025)
di: Frikha, Ahmed, et al.
Pubblicazione: (2025)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
di: Kurochkin, Vadim, et al.
Pubblicazione: (2025)
di: Kurochkin, Vadim, et al.
Pubblicazione: (2025)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
di: Karvonen, Adam, et al.
Pubblicazione: (2024)
di: Karvonen, Adam, et al.
Pubblicazione: (2024)
MoRFI: Monotonic Sparse Autoencoder Feature Identification
di: Dimakopoulos, Dimitris, et al.
Pubblicazione: (2026)
di: Dimakopoulos, Dimitris, et al.
Pubblicazione: (2026)
On the Robustness of Transformers against Context Hijacking for Linear Classification
di: Li, Tianle, et al.
Pubblicazione: (2025)
di: Li, Tianle, et al.
Pubblicazione: (2025)
GUARD: Guided Unlearning and Retention via Data Attribution for Large Language Models
di: Niu, Peizhi, et al.
Pubblicazione: (2025)
di: Niu, Peizhi, et al.
Pubblicazione: (2025)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
di: Lan, Michael, et al.
Pubblicazione: (2024)
di: Lan, Michael, et al.
Pubblicazione: (2024)
Towards Reasoning-Preserving Unlearning in Multimodal Large Language Models
di: Li, Hongji, et al.
Pubblicazione: (2025)
di: Li, Hongji, et al.
Pubblicazione: (2025)
Sparse Autoencoder Features for Classifications and Transferability
di: Gallifant, Jack, et al.
Pubblicazione: (2025)
di: Gallifant, Jack, et al.
Pubblicazione: (2025)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
di: Yu, Zhuohao, et al.
Pubblicazione: (2025)
di: Yu, Zhuohao, et al.
Pubblicazione: (2025)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
di: Li, Aaron J., et al.
Pubblicazione: (2025)
di: Li, Aaron J., et al.
Pubblicazione: (2025)
Large Language Model Unlearning via Embedding-Corrupted Prompts
di: Liu, Chris Yuhao, et al.
Pubblicazione: (2024)
di: Liu, Chris Yuhao, et al.
Pubblicazione: (2024)
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
di: Weng, Jiaqi, et al.
Pubblicazione: (2025)
di: Weng, Jiaqi, et al.
Pubblicazione: (2025)
A Neuro-inspired Interpretation of Unlearning in Large Language Models through Sample-level Unlearning Difficulty
di: Feng, Xiaohua, et al.
Pubblicazione: (2025)
di: Feng, Xiaohua, et al.
Pubblicazione: (2025)
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
di: Guo, Phillip, et al.
Pubblicazione: (2024)
di: Guo, Phillip, et al.
Pubblicazione: (2024)
DRESSing Up LLM: Efficient Stylized Question-Answering via Style Subspace Editing
di: Ma, Xinyu, et al.
Pubblicazione: (2025)
di: Ma, Xinyu, et al.
Pubblicazione: (2025)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning
di: Bi, Jiaxi, et al.
Pubblicazione: (2026)
di: Bi, Jiaxi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2025) -
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
di: Wang, Xu, et al.
Pubblicazione: (2025) -
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2026) -
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
di: Yamashita, Tomoya, et al.
Pubblicazione: (2025) -
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
di: Jing, Yi, et al.
Pubblicazione: (2026)