From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Chae, Kyubyung, Jin, Hyunbin, Kim, Taesup |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
What If TSF: A Benchmark for Reframing Forecasting as Scenario-Guided Multimodal Forecasting
by: Jang, Jinkwan, et al.
Published: (2026)
by: Jang, Jinkwan, et al.
Published: (2026)
Goal-guided Generative Prompt Injection Attack on Large Language Models
by: Zhang, Chong, et al.
Published: (2024)
by: Zhang, Chong, et al.
Published: (2024)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
by: Wang, Haoran, et al.
Published: (2023)
by: Wang, Haoran, et al.
Published: (2023)
May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks
by: Pandya, Nishit V., et al.
Published: (2025)
by: Pandya, Nishit V., et al.
Published: (2025)
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
by: Fujinuma, Yoshinari, et al.
Published: (2026)
by: Fujinuma, Yoshinari, et al.
Published: (2026)
What Matters For Safety Alignment?
by: Li, Xing, et al.
Published: (2026)
by: Li, Xing, et al.
Published: (2026)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
by: Wang, Xilong, et al.
Published: (2026)
by: Wang, Xilong, et al.
Published: (2026)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
by: Li, Jianwei, et al.
Published: (2025)
by: Li, Jianwei, et al.
Published: (2025)
Beyond the Tip of Efficiency: Uncovering the Submerged Threats of Jailbreak Attacks in Small Language Models
by: Yi, Sibo, et al.
Published: (2025)
by: Yi, Sibo, et al.
Published: (2025)
Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models
by: Belkhiter, Yannis, et al.
Published: (2026)
by: Belkhiter, Yannis, et al.
Published: (2026)
Claim-Guided Textual Backdoor Attack for Practical Applications
by: Song, Minkyoo, et al.
Published: (2024)
by: Song, Minkyoo, et al.
Published: (2024)
A Content-Based Framework for Cybersecurity Refusal Decisions in Large Language Models
by: Linder, Noa, et al.
Published: (2026)
by: Linder, Noa, et al.
Published: (2026)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
by: Hu, Xiaomeng, et al.
Published: (2024)
by: Hu, Xiaomeng, et al.
Published: (2024)
IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
by: An, Hengyu, et al.
Published: (2025)
by: An, Hengyu, et al.
Published: (2025)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
by: Hu, Xulin, et al.
Published: (2026)
by: Hu, Xulin, et al.
Published: (2026)
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
by: Pisano, Matthew, et al.
Published: (2023)
by: Pisano, Matthew, et al.
Published: (2023)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
by: Liu, Yupei, et al.
Published: (2023)
by: Liu, Yupei, et al.
Published: (2023)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
by: Kumar, Anurakt, et al.
Published: (2024)
by: Kumar, Anurakt, et al.
Published: (2024)
An Investigation on Group Query Hallucination Attacks
by: Miao, Kehao, et al.
Published: (2025)
by: Miao, Kehao, et al.
Published: (2025)
Superficial Safety Alignment Hypothesis
by: Li, Jianwei, et al.
Published: (2024)
by: Li, Jianwei, et al.
Published: (2024)
Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection
by: Liang, Zi, et al.
Published: (2026)
by: Liang, Zi, et al.
Published: (2026)
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content
by: Guo, Ruoqi, et al.
Published: (2026)
by: Guo, Ruoqi, et al.
Published: (2026)
Safety Alignment Should Be Made More Than Just A Few Attention Heads
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment
by: Wang, Kun, et al.
Published: (2026)
by: Wang, Kun, et al.
Published: (2026)
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
by: Nawal, Aditya, et al.
Published: (2026)
by: Nawal, Aditya, et al.
Published: (2026)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
by: Campbell, David, et al.
Published: (2026)
by: Campbell, David, et al.
Published: (2026)
From Threat Intelligence to Firewall Rules: Semantic Relations in Hybrid AI Agent and Expert System Architectures
by: Bonfanti, Chiara, et al.
Published: (2026)
by: Bonfanti, Chiara, et al.
Published: (2026)
AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models
by: Zhang, Jinchuan, et al.
Published: (2025)
by: Zhang, Jinchuan, et al.
Published: (2025)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
by: Wang, Yanbo, et al.
Published: (2026)
by: Wang, Yanbo, et al.
Published: (2026)
Prompt Injection as Role Confusion
by: Ye, Charles, et al.
Published: (2026)
by: Ye, Charles, et al.
Published: (2026)
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
by: Li, Tianhao, et al.
Published: (2024)
by: Li, Tianhao, et al.
Published: (2024)
Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
by: Yin, Qingyu, et al.
Published: (2025)
by: Yin, Qingyu, et al.
Published: (2025)
Resource Consumption Threats in Large Language Models
by: Zhang, Yuanhe, et al.
Published: (2026)
by: Zhang, Yuanhe, et al.
Published: (2026)
The Ethics of Interaction: Mitigating Security Threats in LLMs
by: Kumar, Ashutosh, et al.
Published: (2024)
by: Kumar, Ashutosh, et al.
Published: (2024)
DMFI: A Dual-Modality Log Analysis Framework for Insider Threat Detection with LoRA-Tuned Language Models
by: Kong, Kaichuan, et al.
Published: (2025)
by: Kong, Kaichuan, et al.
Published: (2025)
Mitigating Hallucination in Abstractive Summarization with Domain-Conditional Mutual Information
by: Chae, Kyubyung, et al.
Published: (2024)
by: Chae, Kyubyung, et al.
Published: (2024)
Similar Items
-
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
by: Li, Hao, et al.
Published: (2026) -
What If TSF: A Benchmark for Reframing Forecasting as Scenario-Guided Multimodal Forecasting
by: Jang, Jinkwan, et al.
Published: (2026) -
Goal-guided Generative Prompt Injection Attack on Large Language Models
by: Zhang, Chong, et al.
Published: (2024) -
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024) -
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
by: Xu, Zhangchen, et al.
Published: (2024)