Steering Language Model Refusal with Sparse Autoencoders
Fuente:
arXiv
Salvato in:
| Autori principali: | O'Brien, Kyle, Majercak, David, Fernandes, Xavier, Edgar, Richard, Bullwinkel, Blake, Chen, Jingya, Nori, Harsha, Carignan, Dean, Horvitz, Eric, Poursabzi-Sangdeh, Forough |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond
di: Nori, Harsha, et al.
Pubblicazione: (2024)
di: Nori, Harsha, et al.
Pubblicazione: (2024)
Understanding Refusal in Language Models with Sparse Autoencoders
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025)
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025)
Sparse identification of nonlinear dynamics in the presence of library and system uncertainty
di: O'Brien, Andrew
Pubblicazione: (2024)
di: O'Brien, Andrew
Pubblicazione: (2024)
Can Revealed Preferences Clarify LLM Alignment and Steering?
di: Yamin, Khurram, et al.
Pubblicazione: (2026)
di: Yamin, Khurram, et al.
Pubblicazione: (2026)
Media Integrity and Authentication: Status, Directions, and Futures
di: Young, Jessica, et al.
Pubblicazione: (2026)
di: Young, Jessica, et al.
Pubblicazione: (2026)
Beyond the Single Turn: Reframing Refusals as Dynamic Experiences Embedded in the Context of Mental Health Support Interactions with LLMs
di: Tang, Ningjing, et al.
Pubblicazione: (2026)
di: Tang, Ningjing, et al.
Pubblicazione: (2026)
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
di: DeLeeuw, Caleb
Pubblicazione: (2026)
di: DeLeeuw, Caleb
Pubblicazione: (2026)
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
di: Chou, Cheng-Ting, et al.
Pubblicazione: (2025)
di: Chou, Cheng-Ting, et al.
Pubblicazione: (2025)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
di: Alagharu, Rishab, et al.
Pubblicazione: (2026)
di: Alagharu, Rishab, et al.
Pubblicazione: (2026)
Improving Instruction-Following in Language Models through Activation Steering
di: Stolfo, Alessandro, et al.
Pubblicazione: (2024)
di: Stolfo, Alessandro, et al.
Pubblicazione: (2024)
JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models
di: Geng, Saibo, et al.
Pubblicazione: (2025)
di: Geng, Saibo, et al.
Pubblicazione: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)
di: Chalnev, Sviatoslav, et al.
Pubblicazione: (2024)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
Graph-Regularized Sparse Autoencoders for LLM Safety Steering
di: Yeon, Jehyeok, et al.
Pubblicazione: (2025)
di: Yeon, Jehyeok, et al.
Pubblicazione: (2025)
Programming Refusal with Conditional Activation Steering
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
di: Ghosh, Shaona, et al.
Pubblicazione: (2025)
di: Ghosh, Shaona, et al.
Pubblicazione: (2025)
Florida's East Coast Inlets: shoreline effects and recommended action
di: Dean, Robert G, et al.
Pubblicazione: (1987)
di: Dean, Robert G, et al.
Pubblicazione: (1987)
Florida's West Coast inlets: shoreline effects and recommended action
di: Dean, Robert G., et al.
Pubblicazione: (1987)
di: Dean, Robert G., et al.
Pubblicazione: (1987)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
di: Cho, Seonglae, et al.
Pubblicazione: (2025)
di: Cho, Seonglae, et al.
Pubblicazione: (2025)
SteerRM: Debiasing Reward Models via Sparse Autoencoders
di: Sun, Mengyuan, et al.
Pubblicazione: (2026)
di: Sun, Mengyuan, et al.
Pubblicazione: (2026)
Interpreting and Steering Protein Language Models through Sparse Autoencoders
di: Garcia, Edith Natalia Villegas, et al.
Pubblicazione: (2025)
di: Garcia, Edith Natalia Villegas, et al.
Pubblicazione: (2025)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
di: Yu, Zhuohao, et al.
Pubblicazione: (2025)
di: Yu, Zhuohao, et al.
Pubblicazione: (2025)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
di: Fang, Yi, et al.
Pubblicazione: (2026)
di: Fang, Yi, et al.
Pubblicazione: (2026)
SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs
di: Härle, Ruben, et al.
Pubblicazione: (2024)
di: Härle, Ruben, et al.
Pubblicazione: (2024)
AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
di: Sheng, Leheng, et al.
Pubblicazione: (2025)
di: Sheng, Leheng, et al.
Pubblicazione: (2025)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
di: García-Ferrero, Iker, et al.
Pubblicazione: (2025)
di: García-Ferrero, Iker, et al.
Pubblicazione: (2025)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
di: Yang, Jiaxi, et al.
Pubblicazione: (2026)
di: Yang, Jiaxi, et al.
Pubblicazione: (2026)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
di: Zhao, Haiyan, et al.
Pubblicazione: (2025)
Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
di: Ghussin, Yusser Al, et al.
Pubblicazione: (2026)
di: Ghussin, Yusser Al, et al.
Pubblicazione: (2026)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
di: Jørgensen, Mikkel Godsk, et al.
Pubblicazione: (2026)
di: Jørgensen, Mikkel Godsk, et al.
Pubblicazione: (2026)
Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
di: Wang, Anyi, et al.
Pubblicazione: (2025)
di: Wang, Anyi, et al.
Pubblicazione: (2025)
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
di: Wu, Xuansheng, et al.
Pubblicazione: (2025)
CBMAS: Cognitive Behavioral Modeling via Activation Steering
di: Ismail, Ahmed H., et al.
Pubblicazione: (2026)
di: Ismail, Ahmed H., et al.
Pubblicazione: (2026)
Elephants Never Forget: Testing Language Models for Memorization of Tabular Data
di: Bordt, Sebastian, et al.
Pubblicazione: (2024)
di: Bordt, Sebastian, et al.
Pubblicazione: (2024)
SonamicExamples
di: O'Brien, Harry
Pubblicazione: (2026)
di: O'Brien, Harry
Pubblicazione: (2026)
Contrasting nonstructural carbohydrate dynamics of tropical tree seedlings under water deficit and variability
di: O'Brien, Michael
Pubblicazione: (2026)
di: O'Brien, Michael
Pubblicazione: (2026)
Geant4 Simulated Dataset for the Relativistic Electron and Proton Telescope integrated little experiment-3
di: O'Brien, Declan
Pubblicazione: (2025)
di: O'Brien, Declan
Pubblicazione: (2025)
Developing Pre-Supernova Neutrino Model Support for sntools
di: O'Brien, Ellie
Pubblicazione: (2026)
di: O'Brien, Ellie
Pubblicazione: (2026)
Traversing European Coastlines (TREC) particle count and meterological data from land (2023-2024)
di: O'Brien, James
Pubblicazione: (2026)
di: O'Brien, James
Pubblicazione: (2026)
Documenti analoghi
-
From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond
di: Nori, Harsha, et al.
Pubblicazione: (2024) -
Understanding Refusal in Language Models with Sparse Autoencoders
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025) -
Sparse identification of nonlinear dynamics in the presence of library and system uncertainty
di: O'Brien, Andrew
Pubblicazione: (2024) -
Can Revealed Preferences Clarify LLM Alignment and Steering?
di: Yamin, Khurram, et al.
Pubblicazione: (2026) -
Media Integrity and Authentication: Status, Directions, and Futures
di: Young, Jessica, et al.
Pubblicazione: (2026)