LatentGuard: Controllable Latent Steering for Robust Refusal of Attacks and Reliable Response Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shu, Huizhen, Li, Xuying, Li, Zhuo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Resurgence of GCG Adversarial Attacks on Large Language Models
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)
Latent-space Attacks for Refusal Evasion in Language Models
von: Piras, Giorgio, et al.
Veröffentlicht: (2026)
von: Piras, Giorgio, et al.
Veröffentlicht: (2026)
LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries
von: Ren, Xuancheng, et al.
Veröffentlicht: (2026)
von: Ren, Xuancheng, et al.
Veröffentlicht: (2026)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
von: Alagharu, Rishab, et al.
Veröffentlicht: (2026)
von: Alagharu, Rishab, et al.
Veröffentlicht: (2026)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
von: Hu, Xulin, et al.
Veröffentlicht: (2026)
von: Hu, Xulin, et al.
Veröffentlicht: (2026)
Exploring the Personality Traits of LLMs through Latent Features Steering
von: Yang, Shu, et al.
Veröffentlicht: (2024)
von: Yang, Shu, et al.
Veröffentlicht: (2024)
Unveiling and Steering Connectome Organization with Interpretable Latent Variables
von: Li, Yubin, et al.
Veröffentlicht: (2025)
von: Li, Yubin, et al.
Veröffentlicht: (2025)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
von: García-Ferrero, Iker, et al.
Veröffentlicht: (2025)
von: García-Ferrero, Iker, et al.
Veröffentlicht: (2025)
Targeting the Core: A Simple and Effective Method to Attack RAG-based Agents via Direct LLM Manipulation
von: Li, Xuying, et al.
Veröffentlicht: (2024)
von: Li, Xuying, et al.
Veröffentlicht: (2024)
Steer LLM Latents for Hallucination Detection
von: Park, Seongheon, et al.
Veröffentlicht: (2025)
von: Park, Seongheon, et al.
Veröffentlicht: (2025)
Latent Guard: a Safety Framework for Text-to-image Generation
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
RISER: Orchestrating Latent Reasoning Skills for Adaptive Activation Steering
von: Ye, Wencheng, et al.
Veröffentlicht: (2026)
von: Ye, Wencheng, et al.
Veröffentlicht: (2026)
Programming Refusal with Conditional Activation Steering
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts
von: Bhargav, Samaksh, et al.
Veröffentlicht: (2025)
von: Bhargav, Samaksh, et al.
Veröffentlicht: (2025)
Controllable Mathematical Reasoning via Self-Optimizing Thought Vectors
von: LI, Xuying
Veröffentlicht: (2025)
von: LI, Xuying
Veröffentlicht: (2025)
Latent Space Disentanglement via Activation Steering for Interpretable Attribute Control in Symbolic Music Generation
von: Prokopiou, Ioannis, et al.
Veröffentlicht: (2026)
von: Prokopiou, Ioannis, et al.
Veröffentlicht: (2026)
Transferable Latent-to-Latent Locomotion Policy for Efficient and Versatile Motion Control of Diverse Legged Robots
von: Zheng, Ziang, et al.
Veröffentlicht: (2025)
von: Zheng, Ziang, et al.
Veröffentlicht: (2025)
AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
von: Sheng, Leheng, et al.
Veröffentlicht: (2025)
von: Sheng, Leheng, et al.
Veröffentlicht: (2025)
Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs
von: Li, Jiakang, et al.
Veröffentlicht: (2026)
von: Li, Jiakang, et al.
Veröffentlicht: (2026)
Output Length Effect on DeepSeek-R1's Safety in Forced Thinking
von: Li, Xuying, et al.
Veröffentlicht: (2025)
von: Li, Xuying, et al.
Veröffentlicht: (2025)
Learn to Refuse: Making Large Language Models More Controllable and Reliable through Knowledge Scope Limitation and Refusal Mechanism
von: Cao, Lang
Veröffentlicht: (2023)
von: Cao, Lang
Veröffentlicht: (2023)
Preemptive Detection and Steering of LLM Misalignment via Latent Reachability
von: Karnik, Sathwik, et al.
Veröffentlicht: (2025)
von: Karnik, Sathwik, et al.
Veröffentlicht: (2025)
On Effects of Steering Latent Representation for Large Language Model Unlearning
von: Huu-Tien, Dang, et al.
Veröffentlicht: (2024)
von: Huu-Tien, Dang, et al.
Veröffentlicht: (2024)
Spatial-Aware Latent Initialization for Controllable Image Generation
von: Sun, Wenqiang, et al.
Veröffentlicht: (2024)
von: Sun, Wenqiang, et al.
Veröffentlicht: (2024)
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)
Latent Action Control for Reasoning-Guided Unified Image Generation
von: Zhai, Fuxiang, et al.
Veröffentlicht: (2026)
von: Zhai, Fuxiang, et al.
Veröffentlicht: (2026)
Controllable and Stealthy Shilling Attacks via Dispersive Latent Diffusion
von: Qiao, Shutong, et al.
Veröffentlicht: (2025)
von: Qiao, Shutong, et al.
Veröffentlicht: (2025)
Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal
von: Yang, Kia-Jüng, et al.
Veröffentlicht: (2026)
von: Yang, Kia-Jüng, et al.
Veröffentlicht: (2026)
Latent Policy Steering with Embodiment-Agnostic Pretrained World Models
von: Wang, Yiqi, et al.
Veröffentlicht: (2025)
von: Wang, Yiqi, et al.
Veröffentlicht: (2025)
Learning Latent Dynamic Robust Representations for World Models
von: Sun, Ruixiang, et al.
Veröffentlicht: (2024)
von: Sun, Ruixiang, et al.
Veröffentlicht: (2024)
Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
von: Egbuna, Nathan, et al.
Veröffentlicht: (2025)
von: Egbuna, Nathan, et al.
Veröffentlicht: (2025)
Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs
von: Liu, Andy Zeyi, et al.
Veröffentlicht: (2026)
von: Liu, Andy Zeyi, et al.
Veröffentlicht: (2026)
FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding
von: Yang, Jinghan, et al.
Veröffentlicht: (2026)
von: Yang, Jinghan, et al.
Veröffentlicht: (2026)
In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
von: Liu, Sheng, et al.
Veröffentlicht: (2023)
von: Liu, Sheng, et al.
Veröffentlicht: (2023)
Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented Generation
von: Zhang, Wentao, et al.
Veröffentlicht: (2026)
von: Zhang, Wentao, et al.
Veröffentlicht: (2026)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
von: Yang, Jiaxi, et al.
Veröffentlicht: (2026)
von: Yang, Jiaxi, et al.
Veröffentlicht: (2026)
Precision Knowledge Editing: Enhancing Safety in Large Language Models
von: Li, Xuying, et al.
Veröffentlicht: (2024)
von: Li, Xuying, et al.
Veröffentlicht: (2024)
Robust Latent Matters: Boosting Image Generation with Sampling Error Synthesis
von: Qiu, Kai, et al.
Veröffentlicht: (2025)
von: Qiu, Kai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Resurgence of GCG Adversarial Attacks on Large Language Models
von: Tan, Yuting, et al.
Veröffentlicht: (2025) -
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
von: Shu, Huizhen, et al.
Veröffentlicht: (2025) -
Latent-space Attacks for Refusal Evasion in Language Models
von: Piras, Giorgio, et al.
Veröffentlicht: (2026) -
LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries
von: Ren, Xuancheng, et al.
Veröffentlicht: (2026) -
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
von: Alagharu, Rishab, et al.
Veröffentlicht: (2026)