Steering Language Model Refusal with Sparse Autoencoders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | O'Brien, Kyle, Majercak, David, Fernandes, Xavier, Edgar, Richard, Bullwinkel, Blake, Chen, Jingya, Nori, Harsha, Carignan, Dean, Horvitz, Eric, Poursabzi-Sangdeh, Forough |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond
von: Nori, Harsha, et al.
Veröffentlicht: (2024)
von: Nori, Harsha, et al.
Veröffentlicht: (2024)
Understanding Refusal in Language Models with Sparse Autoencoders
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
Sparse identification of nonlinear dynamics in the presence of library and system uncertainty
von: O'Brien, Andrew
Veröffentlicht: (2024)
von: O'Brien, Andrew
Veröffentlicht: (2024)
Can Revealed Preferences Clarify LLM Alignment and Steering?
von: Yamin, Khurram, et al.
Veröffentlicht: (2026)
von: Yamin, Khurram, et al.
Veröffentlicht: (2026)
Media Integrity and Authentication: Status, Directions, and Futures
von: Young, Jessica, et al.
Veröffentlicht: (2026)
von: Young, Jessica, et al.
Veröffentlicht: (2026)
Beyond the Single Turn: Reframing Refusals as Dynamic Experiences Embedded in the Context of Mental Health Support Interactions with LLMs
von: Tang, Ningjing, et al.
Veröffentlicht: (2026)
von: Tang, Ningjing, et al.
Veröffentlicht: (2026)
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
von: DeLeeuw, Caleb
Veröffentlicht: (2026)
von: DeLeeuw, Caleb
Veröffentlicht: (2026)
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
von: Chou, Cheng-Ting, et al.
Veröffentlicht: (2025)
von: Chou, Cheng-Ting, et al.
Veröffentlicht: (2025)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
von: Alagharu, Rishab, et al.
Veröffentlicht: (2026)
von: Alagharu, Rishab, et al.
Veröffentlicht: (2026)
Improving Instruction-Following in Language Models through Activation Steering
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models
von: Geng, Saibo, et al.
Veröffentlicht: (2025)
von: Geng, Saibo, et al.
Veröffentlicht: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
von: Hua, Zhenglin, et al.
Veröffentlicht: (2025)
von: Hua, Zhenglin, et al.
Veröffentlicht: (2025)
Graph-Regularized Sparse Autoencoders for LLM Safety Steering
von: Yeon, Jehyeok, et al.
Veröffentlicht: (2025)
von: Yeon, Jehyeok, et al.
Veröffentlicht: (2025)
Programming Refusal with Conditional Activation Steering
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
von: Ghosh, Shaona, et al.
Veröffentlicht: (2025)
von: Ghosh, Shaona, et al.
Veröffentlicht: (2025)
Florida's East Coast Inlets: shoreline effects and recommended action
von: Dean, Robert G, et al.
Veröffentlicht: (1987)
von: Dean, Robert G, et al.
Veröffentlicht: (1987)
Florida's West Coast inlets: shoreline effects and recommended action
von: Dean, Robert G., et al.
Veröffentlicht: (1987)
von: Dean, Robert G., et al.
Veröffentlicht: (1987)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
SteerRM: Debiasing Reward Models via Sparse Autoencoders
von: Sun, Mengyuan, et al.
Veröffentlicht: (2026)
von: Sun, Mengyuan, et al.
Veröffentlicht: (2026)
Interpreting and Steering Protein Language Models through Sparse Autoencoders
von: Garcia, Edith Natalia Villegas, et al.
Veröffentlicht: (2025)
von: Garcia, Edith Natalia Villegas, et al.
Veröffentlicht: (2025)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2025)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
von: Fang, Yi, et al.
Veröffentlicht: (2026)
von: Fang, Yi, et al.
Veröffentlicht: (2026)
SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs
von: Härle, Ruben, et al.
Veröffentlicht: (2024)
von: Härle, Ruben, et al.
Veröffentlicht: (2024)
AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
von: Sheng, Leheng, et al.
Veröffentlicht: (2025)
von: Sheng, Leheng, et al.
Veröffentlicht: (2025)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
von: García-Ferrero, Iker, et al.
Veröffentlicht: (2025)
von: García-Ferrero, Iker, et al.
Veröffentlicht: (2025)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
von: Yang, Jiaxi, et al.
Veröffentlicht: (2026)
von: Yang, Jiaxi, et al.
Veröffentlicht: (2026)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
von: Ghussin, Yusser Al, et al.
Veröffentlicht: (2026)
von: Ghussin, Yusser Al, et al.
Veröffentlicht: (2026)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
von: Jørgensen, Mikkel Godsk, et al.
Veröffentlicht: (2026)
von: Jørgensen, Mikkel Godsk, et al.
Veröffentlicht: (2026)
Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
von: Wang, Anyi, et al.
Veröffentlicht: (2025)
von: Wang, Anyi, et al.
Veröffentlicht: (2025)
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
von: Wu, Xuansheng, et al.
Veröffentlicht: (2025)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2025)
CBMAS: Cognitive Behavioral Modeling via Activation Steering
von: Ismail, Ahmed H., et al.
Veröffentlicht: (2026)
von: Ismail, Ahmed H., et al.
Veröffentlicht: (2026)
Elephants Never Forget: Testing Language Models for Memorization of Tabular Data
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
SonamicExamples
von: O'Brien, Harry
Veröffentlicht: (2026)
von: O'Brien, Harry
Veröffentlicht: (2026)
Contrasting nonstructural carbohydrate dynamics of tropical tree seedlings under water deficit and variability
von: O'Brien, Michael
Veröffentlicht: (2026)
von: O'Brien, Michael
Veröffentlicht: (2026)
Geant4 Simulated Dataset for the Relativistic Electron and Proton Telescope integrated little experiment-3
von: O'Brien, Declan
Veröffentlicht: (2025)
von: O'Brien, Declan
Veröffentlicht: (2025)
Developing Pre-Supernova Neutrino Model Support for sntools
von: O'Brien, Ellie
Veröffentlicht: (2026)
von: O'Brien, Ellie
Veröffentlicht: (2026)
Traversing European Coastlines (TREC) particle count and meterological data from land (2023-2024)
von: O'Brien, James
Veröffentlicht: (2026)
von: O'Brien, James
Veröffentlicht: (2026)
Ähnliche Einträge
-
From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond
von: Nori, Harsha, et al.
Veröffentlicht: (2024) -
Understanding Refusal in Language Models with Sparse Autoencoders
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025) -
Sparse identification of nonlinear dynamics in the presence of library and system uncertainty
von: O'Brien, Andrew
Veröffentlicht: (2024) -
Can Revealed Preferences Clarify LLM Alignment and Steering?
von: Yamin, Khurram, et al.
Veröffentlicht: (2026) -
Media Integrity and Authentication: Status, Directions, and Futures
von: Young, Jessica, et al.
Veröffentlicht: (2026)