SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | He, Zirui, Zhao, Haiyan, Qiao, Yiran, Yang, Fan, Payani, Ali, Ma, Jing, Du, Mengnan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models
di: He, Zirui, et al.
Pubblicazione: (2025)
di: He, Zirui, et al.
Pubblicazione: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation
di: Zhao, Haiyan, et al.
Pubblicazione: (2026)
di: Zhao, Haiyan, et al.
Pubblicazione: (2026)
LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories
di: He, Zirui, et al.
Pubblicazione: (2025)
di: He, Zirui, et al.
Pubblicazione: (2025)
Rep2Text: Decoding Full Text from a Single LLM Token Representation
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution
di: Zhao, Haiyan, et al.
Pubblicazione: (2024)
di: Zhao, Haiyan, et al.
Pubblicazione: (2024)
Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
di: Li, Zihao, et al.
Pubblicazione: (2024)
di: Li, Zihao, et al.
Pubblicazione: (2024)
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection
di: Zhang, Huopu, et al.
Pubblicazione: (2025)
di: Zhang, Huopu, et al.
Pubblicazione: (2025)
Towards Uncovering How Large Language Model Works: An Explainability Perspective
di: Zhao, Haiyan, et al.
Pubblicazione: (2024)
di: Zhao, Haiyan, et al.
Pubblicazione: (2024)
Improving LLM Reasoning through Interpretable Role-Playing Steering
di: Wang, Anyi, et al.
Pubblicazione: (2025)
di: Wang, Anyi, et al.
Pubblicazione: (2025)
NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models
di: Liu, Weiqi, et al.
Pubblicazione: (2026)
di: Liu, Weiqi, et al.
Pubblicazione: (2026)
Exploring Multilingual Probing in Large Language Models: A Cross-Language Analysis
di: Li, Daoyang, et al.
Pubblicazione: (2024)
di: Li, Daoyang, et al.
Pubblicazione: (2024)
SAGE: An Agentic Explainer Framework for Interpreting SAE Features in Language Models
di: Han, Jiaojiao, et al.
Pubblicazione: (2025)
di: Han, Jiaojiao, et al.
Pubblicazione: (2025)
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations
di: Yao, Yifei, et al.
Pubblicazione: (2025)
di: Yao, Yifei, et al.
Pubblicazione: (2025)
A Comparative Analysis of Sparse Autoencoder and Activation Difference in Language Model Steering
di: Xie, Jiaqing
Pubblicazione: (2025)
di: Xie, Jiaqing
Pubblicazione: (2025)
Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models
di: Li, Zihao, et al.
Pubblicazione: (2025)
di: Li, Zihao, et al.
Pubblicazione: (2025)
Fine-Grained Interpretation of Political Opinions in Large Language Models
di: Hu, Jingyu, et al.
Pubblicazione: (2025)
di: Hu, Jingyu, et al.
Pubblicazione: (2025)
SteerRM: Debiasing Reward Models via Sparse Autoencoders
di: Sun, Mengyuan, et al.
Pubblicazione: (2026)
di: Sun, Mengyuan, et al.
Pubblicazione: (2026)
FinAnchor: Aligned Multi-Model Representations for Financial Prediction
di: He, Zirui, et al.
Pubblicazione: (2026)
di: He, Zirui, et al.
Pubblicazione: (2026)
Improving Instruction-Following in Language Models through Activation Steering
di: Stolfo, Alessandro, et al.
Pubblicazione: (2024)
di: Stolfo, Alessandro, et al.
Pubblicazione: (2024)
Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech
di: Du, Hongfei, et al.
Pubblicazione: (2026)
di: Du, Hongfei, et al.
Pubblicazione: (2026)
Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model
di: Xiong, Siheng, et al.
Pubblicazione: (2024)
di: Xiong, Siheng, et al.
Pubblicazione: (2024)
Activation Scaling for Steering and Interpreting Language Models
di: Stoehr, Niklas, et al.
Pubblicazione: (2024)
di: Stoehr, Niklas, et al.
Pubblicazione: (2024)
Counterfactual Visual Explanation via Causally-Guided Adversarial Steering
di: Qiao, Yiran, et al.
Pubblicazione: (2025)
di: Qiao, Yiran, et al.
Pubblicazione: (2025)
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2026)
di: Wang, Xu, et al.
Pubblicazione: (2026)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
Interpretable LLM Guardrails via Sparse Representation Steering
di: He, Zeqing, et al.
Pubblicazione: (2025)
di: He, Zeqing, et al.
Pubblicazione: (2025)
SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs
di: Härle, Ruben, et al.
Pubblicazione: (2024)
di: Härle, Ruben, et al.
Pubblicazione: (2024)
SAIF: Sparse Adversarial and Imperceptible Attack Framework
di: Imtiaz, Tooba, et al.
Pubblicazione: (2022)
di: Imtiaz, Tooba, et al.
Pubblicazione: (2022)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)
Large Language Models Can Learn Temporal Reasoning
di: Xiong, Siheng, et al.
Pubblicazione: (2024)
di: Xiong, Siheng, et al.
Pubblicazione: (2024)
SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector
di: Lee, Kyeongryul, et al.
Pubblicazione: (2025)
di: Lee, Kyeongryul, et al.
Pubblicazione: (2025)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
di: Fang, Yi, et al.
Pubblicazione: (2026)
di: Fang, Yi, et al.
Pubblicazione: (2026)
Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
di: Ghussin, Yusser Al, et al.
Pubblicazione: (2026)
di: Ghussin, Yusser Al, et al.
Pubblicazione: (2026)
Infusing Hierarchical Guidance into Prompt Tuning: A Parameter-Efficient Framework for Multi-level Implicit Discourse Relation Recognition
di: Zhao, Haodong, et al.
Pubblicazione: (2024)
di: Zhao, Haodong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models
di: He, Zirui, et al.
Pubblicazione: (2025) -
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
di: Zhao, Haiyan, et al.
Pubblicazione: (2025) -
Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation
di: Zhao, Haiyan, et al.
Pubblicazione: (2026) -
LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories
di: He, Zirui, et al.
Pubblicazione: (2025) -
Rep2Text: Decoding Full Text from a Single LLM Token Representation
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)