Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Xu, Hu, Yan, Wang, Benyou, Zou, Difan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
di: Wang, Xu, et al.
Pubblicazione: (2025)
di: Wang, Xu, et al.
Pubblicazione: (2025)
Model Unlearning via Sparse Autoencoder Subspace Guided Projections
di: Wang, Xu, et al.
Pubblicazione: (2025)
di: Wang, Xu, et al.
Pubblicazione: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2026)
di: Wang, Xu, et al.
Pubblicazione: (2026)
A Human-Like Reasoning Framework for Multi-Phases Planning Task with Large Language Models
di: Xie, Chengxing, et al.
Pubblicazione: (2024)
di: Xie, Chengxing, et al.
Pubblicazione: (2024)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
di: Jing, Yi, et al.
Pubblicazione: (2026)
di: Jing, Yi, et al.
Pubblicazione: (2026)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
di: Yu, Zhuohao, et al.
Pubblicazione: (2025)
di: Yu, Zhuohao, et al.
Pubblicazione: (2025)
Sparse Autoencoder Features for Classifications and Transferability
di: Gallifant, Jack, et al.
Pubblicazione: (2025)
di: Gallifant, Jack, et al.
Pubblicazione: (2025)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
di: Muchane, Mark, et al.
Pubblicazione: (2025)
di: Muchane, Mark, et al.
Pubblicazione: (2025)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2025)
di: Chanin, David, et al.
Pubblicazione: (2025)
Online Rubrics Elicitation from Pairwise Comparisons
di: Rezaei, MohammadHossein, et al.
Pubblicazione: (2025)
di: Rezaei, MohammadHossein, et al.
Pubblicazione: (2025)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
di: Farnik, Lucy, et al.
Pubblicazione: (2025)
di: Farnik, Lucy, et al.
Pubblicazione: (2025)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
di: Li, Aaron J., et al.
Pubblicazione: (2025)
di: Li, Aaron J., et al.
Pubblicazione: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)
Boosting Protein Language Models with Negative Sample Mining
di: Xu, Yaoyao, et al.
Pubblicazione: (2024)
di: Xu, Yaoyao, et al.
Pubblicazione: (2024)
On the Robustness of Transformers against Context Hijacking for Linear Classification
di: Li, Tianle, et al.
Pubblicazione: (2025)
di: Li, Tianle, et al.
Pubblicazione: (2025)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
di: Zhu, Xudong, et al.
Pubblicazione: (2025)
di: Zhu, Xudong, et al.
Pubblicazione: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2025)
di: Chanin, David, et al.
Pubblicazione: (2025)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
di: Tang, Zhengyang, et al.
Pubblicazione: (2024)
di: Tang, Zhengyang, et al.
Pubblicazione: (2024)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
di: Jørgensen, Mikkel Godsk, et al.
Pubblicazione: (2026)
di: Jørgensen, Mikkel Godsk, et al.
Pubblicazione: (2026)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
di: He, Zirui, et al.
Pubblicazione: (2025)
di: He, Zirui, et al.
Pubblicazione: (2025)
In-context Autoencoder for Context Compression in a Large Language Model
di: Ge, Tao, et al.
Pubblicazione: (2023)
di: Ge, Tao, et al.
Pubblicazione: (2023)
Diversity-driven Data Selection for Language Model Tuning through Sparse Autoencoder
di: Yang, Xianjun, et al.
Pubblicazione: (2025)
di: Yang, Xianjun, et al.
Pubblicazione: (2025)
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
di: Joshi, Shruti, et al.
Pubblicazione: (2025)
di: Joshi, Shruti, et al.
Pubblicazione: (2025)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
di: Lieberum, Tom, et al.
Pubblicazione: (2024)
di: Lieberum, Tom, et al.
Pubblicazione: (2024)
Towards Understanding the Robustness of Sparse Autoencoders
di: Saiyed, Ahson, et al.
Pubblicazione: (2026)
di: Saiyed, Ahson, et al.
Pubblicazione: (2026)
MAPLE: Micro Analysis of Pairwise Language Evolution for Few-Shot Claim Verification
di: Zeng, Xia, et al.
Pubblicazione: (2024)
di: Zeng, Xia, et al.
Pubblicazione: (2024)
Ask more, know better: Reinforce-Learned Prompt Questions for Decision Making with Large Language Models
di: Yan, Xue, et al.
Pubblicazione: (2023)
di: Yan, Xue, et al.
Pubblicazione: (2023)
Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training
di: Li, T. Ed, et al.
Pubblicazione: (2025)
di: Li, T. Ed, et al.
Pubblicazione: (2025)
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
di: Muhamed, Aashiq, et al.
Pubblicazione: (2024)
di: Muhamed, Aashiq, et al.
Pubblicazione: (2024)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
di: Lan, Michael, et al.
Pubblicazione: (2024)
di: Lan, Michael, et al.
Pubblicazione: (2024)
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
di: Weng, Jiaqi, et al.
Pubblicazione: (2025)
di: Weng, Jiaqi, et al.
Pubblicazione: (2025)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
di: Cho, Seonglae, et al.
Pubblicazione: (2025)
di: Cho, Seonglae, et al.
Pubblicazione: (2025)
Dissecting Chronos: Sparse Autoencoders Reveal Causal Feature Hierarchies in Time Series Foundation Models
di: Mishra, Anurag
Pubblicazione: (2026)
di: Mishra, Anurag
Pubblicazione: (2026)
Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts
di: Peng, Kenny, et al.
Pubblicazione: (2025)
di: Peng, Kenny, et al.
Pubblicazione: (2025)
Sparse Autoencoder Decomposition of Clinical Sequence Model Representations: Feature Complexity, Task Specialisation, and Mortality Prediction
di: Sainsbury, Chris, et al.
Pubblicazione: (2026)
di: Sainsbury, Chris, et al.
Pubblicazione: (2026)
CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis
di: Chen, Junying, et al.
Pubblicazione: (2024)
di: Chen, Junying, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
di: Wang, Xu, et al.
Pubblicazione: (2025) -
Model Unlearning via Sparse Autoencoder Subspace Guided Projections
di: Wang, Xu, et al.
Pubblicazione: (2025) -
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2026) -
A Human-Like Reasoning Framework for Multi-Phases Planning Task with Large Language Models
di: Xie, Chengxing, et al.
Pubblicazione: (2024) -
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
di: Jing, Yi, et al.
Pubblicazione: (2026)