Gradient-Controlled Decoding: A Safety Guardrail for LLMs with Dual-Anchor Steering
Fuente:
arXiv
Saved in:
| Main Authors: | Chiniya, Purva, Scaria, Kevin, Chaturvedi, Sagar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Embarrassingly Simple Unsupervised Aspect Based Sentiment Tuple Extraction
by: Scaria, Kevin, et al.
Published: (2024)
by: Scaria, Kevin, et al.
Published: (2024)
Bypassing Safety Guardrails in LLMs Using Humor
by: Cisneros-Velarde, Pedro
Published: (2025)
by: Cisneros-Velarde, Pedro
Published: (2025)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
Guardrail Baselines for Unlearning in LLMs
by: Thaker, Pratiksha, et al.
Published: (2024)
by: Thaker, Pratiksha, et al.
Published: (2024)
Interpretable LLM Guardrails via Sparse Representation Steering
by: He, Zeqing, et al.
Published: (2025)
by: He, Zeqing, et al.
Published: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
by: Siu, Vincent, et al.
Published: (2025)
by: Siu, Vincent, et al.
Published: (2025)
A Lightweight Explainable Guardrail for Prompt Safety
by: Islam, Md Asiful, et al.
Published: (2026)
by: Islam, Md Asiful, et al.
Published: (2026)
MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
by: Yang, Yahan, et al.
Published: (2025)
by: Yang, Yahan, et al.
Published: (2025)
Steering Multimodal Large Language Models Decoding for Context-Aware Safety
by: Liu, Zheyuan, et al.
Published: (2025)
by: Liu, Zheyuan, et al.
Published: (2025)
Lightweight Safety Guardrails Using Fine-tuned BERT Embeddings
by: Zheng, Aaron, et al.
Published: (2024)
by: Zheng, Aaron, et al.
Published: (2024)
Do LLM Agents Mirror Socio-Cognitive Effects in Power-Asymmetric Conversations?
by: Vijjini, Anvesh Rao, et al.
Published: (2026)
by: Vijjini, Anvesh Rao, et al.
Published: (2026)
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
by: Chen, Shuo, et al.
Published: (2025)
by: Chen, Shuo, et al.
Published: (2025)
Capturing Bias Diversity in LLMs
by: Gosavi, Purva Prasad, et al.
Published: (2024)
by: Gosavi, Purva Prasad, et al.
Published: (2024)
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
by: Wang, Yan, et al.
Published: (2026)
by: Wang, Yan, et al.
Published: (2026)
LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
by: An, Heajun, et al.
Published: (2026)
by: An, Heajun, et al.
Published: (2026)
KV Cache Steering for Controlling Frozen LLMs
by: Belitsky, Max, et al.
Published: (2025)
by: Belitsky, Max, et al.
Published: (2025)
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025)
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025)
AV-RIR: Audio-Visual Room Impulse Response Estimation
by: Ratnarajah, Anton, et al.
Published: (2023)
by: Ratnarajah, Anton, et al.
Published: (2023)
Why Do Safety Guardrails Degrade Across Languages?
by: Zhang, Max, et al.
Published: (2026)
by: Zhang, Max, et al.
Published: (2026)
Steer Model beyond Assistant: Controlling System Prompt Strength via Contrastive Decoding
by: Dong, Yijiang River, et al.
Published: (2026)
by: Dong, Yijiang River, et al.
Published: (2026)
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
by: Chu, Hua-Rong, et al.
Published: (2026)
by: Chu, Hua-Rong, et al.
Published: (2026)
Generalization or Memorization: Dynamic Decoding for Mode Steering
by: Zhang, Xuanming
Published: (2025)
by: Zhang, Xuanming
Published: (2025)
Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
by: Huang, Wei-Chieh, et al.
Published: (2025)
by: Huang, Wei-Chieh, et al.
Published: (2025)
Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering
by: Bonaldi, Helena, et al.
Published: (2024)
by: Bonaldi, Helena, et al.
Published: (2024)
GeoSteer: Faithful Chain-of-Thought Steering via Latent Manifold Gradients
by: Kazama, Kentaro, et al.
Published: (2026)
by: Kazama, Kentaro, et al.
Published: (2026)
Cutting Through the Noise: Boosting LLM Performance on Math Word Problems
by: Anantheswaran, Ujjwala, et al.
Published: (2024)
by: Anantheswaran, Ujjwala, et al.
Published: (2024)
Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
by: Ghosh, Shaona, et al.
Published: (2025)
by: Ghosh, Shaona, et al.
Published: (2025)
SteerConf: Steering LLMs for Confidence Elicitation
by: Zhou, Ziang, et al.
Published: (2025)
by: Zhou, Ziang, et al.
Published: (2025)
SGuard-v1: Safety Guardrail for Large Language Models
by: Lee, JoonHo, et al.
Published: (2025)
by: Lee, JoonHo, et al.
Published: (2025)
SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models
by: Lee, Seanie, et al.
Published: (2025)
by: Lee, Seanie, et al.
Published: (2025)
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
by: Nguyen, Duy, et al.
Published: (2025)
by: Nguyen, Duy, et al.
Published: (2025)
FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering
by: Li, Yichen, et al.
Published: (2025)
by: Li, Yichen, et al.
Published: (2025)
CREST: Universal Safety Guardrails Through Cluster-Guided Cross-Lingual Transfer
by: Bansal, Lavish, et al.
Published: (2025)
by: Bansal, Lavish, et al.
Published: (2025)
Guiding Giants: Lightweight Controllers for Weighted Activation Steering in LLMs
by: Hegazy, Amr, et al.
Published: (2025)
by: Hegazy, Amr, et al.
Published: (2025)
Steering LLMs for Culturally Localized Generation
by: Khanuja, Simran, et al.
Published: (2026)
by: Khanuja, Simran, et al.
Published: (2026)
CoSteer: Collaborative Decoding-Time Personalization via Local Delta Steering
by: Lv, Hang, et al.
Published: (2025)
by: Lv, Hang, et al.
Published: (2025)
Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
by: Aswal, Darpan, et al.
Published: (2025)
by: Aswal, Darpan, et al.
Published: (2025)
Steered Generation via Gradient Descent on Sparse Features
by: Bhattacharyya, Sumanta, et al.
Published: (2025)
by: Bhattacharyya, Sumanta, et al.
Published: (2025)
Similar Items
-
Embarrassingly Simple Unsupervised Aspect Based Sentiment Tuple Extraction
by: Scaria, Kevin, et al.
Published: (2024) -
Bypassing Safety Guardrails in LLMs Using Humor
by: Cisneros-Velarde, Pedro
Published: (2025) -
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
by: Ghosh, Sreyan, et al.
Published: (2024) -
Guardrail Baselines for Unlearning in LLMs
by: Thaker, Pratiksha, et al.
Published: (2024) -
Interpretable LLM Guardrails via Sparse Representation Steering
by: He, Zeqing, et al.
Published: (2025)