Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
Fuente:
arXiv
Guardado en:
| Autores principales: | Shu, Dong, Wu, Xuansheng, Zhao, Haiyan, Du, Mengnan, Liu, Ninghao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
por: Shu, Dong, et al.
Publicado: (2025)
por: Shu, Dong, et al.
Publicado: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
por: Zhao, Haiyan, et al.
Publicado: (2025)
por: Zhao, Haiyan, et al.
Publicado: (2025)
Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
por: Wang, Anyi, et al.
Publicado: (2025)
por: Wang, Anyi, et al.
Publicado: (2025)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
por: He, Zirui, et al.
Publicado: (2025)
por: He, Zirui, et al.
Publicado: (2025)
Comparative Analysis of Demonstration Selection Algorithms for LLM In-Context Learning
por: Shu, Dong, et al.
Publicado: (2024)
por: Shu, Dong, et al.
Publicado: (2024)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
por: Wu, Xuansheng, et al.
Publicado: (2023)
por: Wu, Xuansheng, et al.
Publicado: (2023)
Retrieval-enhanced Knowledge Editing in Language Models for Multi-Hop Question Answering
por: Shi, Yucheng, et al.
Publicado: (2024)
por: Shi, Yucheng, et al.
Publicado: (2024)
Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution
por: Zhao, Haiyan, et al.
Publicado: (2024)
por: Zhao, Haiyan, et al.
Publicado: (2024)
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
por: Joshi, Shruti, et al.
Publicado: (2025)
por: Joshi, Shruti, et al.
Publicado: (2025)
Exploring Multilingual Probing in Large Language Models: A Cross-Language Analysis
por: Li, Daoyang, et al.
Publicado: (2024)
por: Li, Daoyang, et al.
Publicado: (2024)
Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
por: Li, Zihao, et al.
Publicado: (2024)
por: Li, Zihao, et al.
Publicado: (2024)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
por: Farnik, Lucy, et al.
Publicado: (2025)
por: Farnik, Lucy, et al.
Publicado: (2025)
Influential Language Data Selection via Gradient Trajectory Pursuit
por: Deng, Zhiwei, et al.
Publicado: (2024)
por: Deng, Zhiwei, et al.
Publicado: (2024)
Rep2Text: Decoding Full Text from a Single LLM Token Representation
por: Zhao, Haiyan, et al.
Publicado: (2025)
por: Zhao, Haiyan, et al.
Publicado: (2025)
Sparse Autoencoder Features for Classifications and Transferability
por: Gallifant, Jack, et al.
Publicado: (2025)
por: Gallifant, Jack, et al.
Publicado: (2025)
Fine-Grained Interpretation of Political Opinions in Large Language Models
por: Hu, Jingyu, et al.
Publicado: (2025)
por: Hu, Jingyu, et al.
Publicado: (2025)
Efficient Long-distance Latent Relation-aware Graph Neural Network for Multi-modal Emotion Recognition in Conversations
por: Shou, Yuntao, et al.
Publicado: (2024)
por: Shou, Yuntao, et al.
Publicado: (2024)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
por: Muchane, Mark, et al.
Publicado: (2025)
por: Muchane, Mark, et al.
Publicado: (2025)
Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
por: Wu, Xuansheng, et al.
Publicado: (2025)
por: Wu, Xuansheng, et al.
Publicado: (2025)
Sparse Autoencoder Decomposition of Clinical Sequence Model Representations: Feature Complexity, Task Specialisation, and Mortality Prediction
por: Sainsbury, Chris, et al.
Publicado: (2026)
por: Sainsbury, Chris, et al.
Publicado: (2026)
Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task Projection
por: Jin, Pengfei, et al.
Publicado: (2024)
por: Jin, Pengfei, et al.
Publicado: (2024)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2025)
por: Chanin, David, et al.
Publicado: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
por: Chalnev, Sviatoslav, et al.
Publicado: (2024)
por: Chalnev, Sviatoslav, et al.
Publicado: (2024)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
por: Li, Aaron J., et al.
Publicado: (2025)
por: Li, Aaron J., et al.
Publicado: (2025)
LESS: Selecting Influential Data for Targeted Instruction Tuning
por: Xia, Mengzhou, et al.
Publicado: (2024)
por: Xia, Mengzhou, et al.
Publicado: (2024)
Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era
por: Wu, Xuansheng, et al.
Publicado: (2024)
por: Wu, Xuansheng, et al.
Publicado: (2024)
Beyond Gradient and Priors in Privacy Attacks: Leveraging Pooler Layer Inputs of Language Models in Federated Learning
por: Li, Jianwei, et al.
Publicado: (2023)
por: Li, Jianwei, et al.
Publicado: (2023)
Diversity-driven Data Selection for Language Model Tuning through Sparse Autoencoder
por: Yang, Xianjun, et al.
Publicado: (2025)
por: Yang, Xianjun, et al.
Publicado: (2025)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
por: Zhu, Xudong, et al.
Publicado: (2025)
por: Zhu, Xudong, et al.
Publicado: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
por: Bhalla, Usha, et al.
Publicado: (2025)
por: Bhalla, Usha, et al.
Publicado: (2025)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2025)
por: Chanin, David, et al.
Publicado: (2025)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
por: Yu, Zhuohao, et al.
Publicado: (2025)
por: Yu, Zhuohao, et al.
Publicado: (2025)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
por: Minegishi, Gouki, et al.
Publicado: (2025)
por: Minegishi, Gouki, et al.
Publicado: (2025)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
por: Jørgensen, Mikkel Godsk, et al.
Publicado: (2026)
por: Jørgensen, Mikkel Godsk, et al.
Publicado: (2026)
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
por: Lee, Gyeong-Geon, et al.
Publicado: (2023)
por: Lee, Gyeong-Geon, et al.
Publicado: (2023)
Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning
por: Li, Xiaochuan, et al.
Publicado: (2024)
por: Li, Xiaochuan, et al.
Publicado: (2024)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
por: Lieberum, Tom, et al.
Publicado: (2024)
por: Lieberum, Tom, et al.
Publicado: (2024)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
por: Wang, Xu, et al.
Publicado: (2026)
por: Wang, Xu, et al.
Publicado: (2026)
Towards Understanding the Robustness of Sparse Autoencoders
por: Saiyed, Ahson, et al.
Publicado: (2026)
por: Saiyed, Ahson, et al.
Publicado: (2026)
BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation
por: Wang, Yun, et al.
Publicado: (2026)
por: Wang, Yun, et al.
Publicado: (2026)
Ejemplares similares
-
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
por: Shu, Dong, et al.
Publicado: (2025) -
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
por: Zhao, Haiyan, et al.
Publicado: (2025) -
Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
por: Wang, Anyi, et al.
Publicado: (2025) -
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
por: He, Zirui, et al.
Publicado: (2025) -
Comparative Analysis of Demonstration Selection Algorithms for LLM In-Context Learning
por: Shu, Dong, et al.
Publicado: (2024)