Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bhargav, Samaksh, Zhu, Zining |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds
von: Wu, Honghan, et al.
Veröffentlicht: (2026)
von: Wu, Honghan, et al.
Veröffentlicht: (2026)
Programming Refusal with Conditional Activation Steering
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation
von: Ajwani, Rohan Deepak, et al.
Veröffentlicht: (2024)
von: Ajwani, Rohan Deepak, et al.
Veröffentlicht: (2024)
AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
von: Sheng, Leheng, et al.
Veröffentlicht: (2025)
von: Sheng, Leheng, et al.
Veröffentlicht: (2025)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
von: Wedgwood, James, et al.
Veröffentlicht: (2026)
von: Wedgwood, James, et al.
Veröffentlicht: (2026)
Distributed Interpretability and Control for Large Language Models
von: Desai, Dev Arpan, et al.
Veröffentlicht: (2026)
von: Desai, Dev Arpan, et al.
Veröffentlicht: (2026)
Dense SAE Latents Are Features, Not Bugs
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
On the Failure of Topic-Matched Contrast Baselines in Multi-Directional Refusal Abliteration
von: Petrov, Valentin
Veröffentlicht: (2026)
von: Petrov, Valentin
Veröffentlicht: (2026)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
von: Yang, Jiaxi, et al.
Veröffentlicht: (2026)
von: Yang, Jiaxi, et al.
Veröffentlicht: (2026)
Positive Unlabeled Contrastive Learning
von: Acharya, Anish, et al.
Veröffentlicht: (2022)
von: Acharya, Anish, et al.
Veröffentlicht: (2022)
Steer Like the LLM: Activation Steering that Mimics Prompting
von: Heyman, Geert, et al.
Veröffentlicht: (2026)
von: Heyman, Geert, et al.
Veröffentlicht: (2026)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
von: Soo, Samuel, et al.
Veröffentlicht: (2025)
What Would You Ask When You First Saw $a^2+b^2=c^2$? Evaluating LLM on Curiosity-Driven Questioning
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2024)
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2024)
Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026)
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026)
Interpreting Affine Recurrence Learning in GPT-style Transformers
von: Bhargav, Samarth, et al.
Veröffentlicht: (2024)
von: Bhargav, Samarth, et al.
Veröffentlicht: (2024)
Salient Information Prompting to Steer Content in Prompt-based Abstractive Summarization
von: Xu, Lei, et al.
Veröffentlicht: (2024)
von: Xu, Lei, et al.
Veröffentlicht: (2024)
Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
von: Zhang, Zeyu, et al.
Veröffentlicht: (2026)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2026)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
Steering Llama 2 via Contrastive Activation Addition
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
ExecTune: Effective Steering of Black-Box LLMs with Guide Models
von: Lingam, Vijay, et al.
Veröffentlicht: (2026)
von: Lingam, Vijay, et al.
Veröffentlicht: (2026)
Efficient Refusal Ablation in LLM through Optimal Transport
von: Nanfack, Geraldin, et al.
Veröffentlicht: (2026)
von: Nanfack, Geraldin, et al.
Veröffentlicht: (2026)
Compositional Literary Primitives in Instruction-Tuned LLMs: Cross-Architectural SAE Features for Self, Style, and Affect
von: Presa, Joao Paulo Cavalcante, et al.
Veröffentlicht: (2026)
von: Presa, Joao Paulo Cavalcante, et al.
Veröffentlicht: (2026)
HelpSteer2-Preference: Complementing Ratings with Preferences
von: Wang, Zhilin, et al.
Veröffentlicht: (2024)
von: Wang, Zhilin, et al.
Veröffentlicht: (2024)
PGCLODA: Prompt-Guided Graph Contrastive Learning for Oligopeptide-Infectious Disease Association Prediction
von: Tan, Dayu, et al.
Veröffentlicht: (2025)
von: Tan, Dayu, et al.
Veröffentlicht: (2025)
ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions
von: Poduval, Prathyush, et al.
Veröffentlicht: (2026)
von: Poduval, Prathyush, et al.
Veröffentlicht: (2026)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
von: Lamb, Tom A., et al.
Veröffentlicht: (2024)
von: Lamb, Tom A., et al.
Veröffentlicht: (2024)
GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
von: Handa, Divij, et al.
Veröffentlicht: (2025)
von: Handa, Divij, et al.
Veröffentlicht: (2025)
CORE: Contrastive Masked Feature Reconstruction on Graphs
von: Bo, Jianyuan, et al.
Veröffentlicht: (2025)
von: Bo, Jianyuan, et al.
Veröffentlicht: (2025)
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models
von: Rahimi, Eliron, et al.
Veröffentlicht: (2026)
von: Rahimi, Eliron, et al.
Veröffentlicht: (2026)
On the Properties of Feature Attribution for Supervised Contrastive Learning
von: Arrighi, Leonardo, et al.
Veröffentlicht: (2026)
von: Arrighi, Leonardo, et al.
Veröffentlicht: (2026)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
von: Alagharu, Rishab, et al.
Veröffentlicht: (2026)
von: Alagharu, Rishab, et al.
Veröffentlicht: (2026)
Angular Steering: Behavior Control via Rotation in Activation Space
von: Vu, Hieu M., et al.
Veröffentlicht: (2025)
von: Vu, Hieu M., et al.
Veröffentlicht: (2025)
Improving Domain Generalization in Contrastive Learning using Adaptive Temperature Control
von: Lewis, Robert, et al.
Veröffentlicht: (2026)
von: Lewis, Robert, et al.
Veröffentlicht: (2026)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
SAEs Are Good for Steering -- If You Select the Right Features
von: Arad, Dana, et al.
Veröffentlicht: (2025)
von: Arad, Dana, et al.
Veröffentlicht: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
Refusing Safe Prompts for Multi-modal Large Language Models
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
Efficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise Replacement
von: Yang, Shu, et al.
Veröffentlicht: (2025)
von: Yang, Shu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds
von: Wu, Honghan, et al.
Veröffentlicht: (2026) -
Programming Refusal with Conditional Activation Steering
von: Lee, Bruce W., et al.
Veröffentlicht: (2024) -
Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation
von: Ajwani, Rohan Deepak, et al.
Veröffentlicht: (2024) -
AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
von: Sheng, Leheng, et al.
Veröffentlicht: (2025) -
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)