One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Arif, Samee, Deng, Naihao, Jin, Zhijing, Mihalcea, Rada |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models
von: Arif, Samee, et al.
Veröffentlicht: (2026)
von: Arif, Samee, et al.
Veröffentlicht: (2026)
Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
von: Deng, Naihao, et al.
Veröffentlicht: (2025)
von: Deng, Naihao, et al.
Veröffentlicht: (2025)
Rethinking Table Instruction Tuning
von: Deng, Naihao, et al.
Veröffentlicht: (2025)
von: Deng, Naihao, et al.
Veröffentlicht: (2025)
One Trigger Token Is Enough: A Defense Strategy for Balancing Safety and Usability in Large Language Models
von: Gu, Haoran, et al.
Veröffentlicht: (2025)
von: Gu, Haoran, et al.
Veröffentlicht: (2025)
Security Attacks on LLM-based Code Completion Tools
von: Cheng, Wen, et al.
Veröffentlicht: (2024)
von: Cheng, Wen, et al.
Veröffentlicht: (2024)
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level
von: Zeng, Xinyi, et al.
Veröffentlicht: (2024)
von: Zeng, Xinyi, et al.
Veröffentlicht: (2024)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
von: Chu, Hua-Rong, et al.
Veröffentlicht: (2026)
von: Chu, Hua-Rong, et al.
Veröffentlicht: (2026)
Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution
von: Zhang, Xiaozhe, et al.
Veröffentlicht: (2026)
von: Zhang, Xiaozhe, et al.
Veröffentlicht: (2026)
The Double-edged Sword of LLM-based Data Reconstruction: Understanding and Mitigating Contextual Vulnerability in Word-level Differential Privacy Text Sanitization
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
von: Singh, Himanshu, et al.
Veröffentlicht: (2026)
von: Singh, Himanshu, et al.
Veröffentlicht: (2026)
DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning
von: Zhang, Junbo, et al.
Veröffentlicht: (2026)
von: Zhang, Junbo, et al.
Veröffentlicht: (2026)
Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak
von: Gu, Haoran, et al.
Veröffentlicht: (2026)
von: Gu, Haoran, et al.
Veröffentlicht: (2026)
The Surprising Universality of LLM Outputs: A Real-Time Verification Primitive
von: Bogdan, Alex, et al.
Veröffentlicht: (2026)
von: Bogdan, Alex, et al.
Veröffentlicht: (2026)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
von: Li, Xirui, et al.
Veröffentlicht: (2024)
von: Li, Xirui, et al.
Veröffentlicht: (2024)
Robust LLM Watermarking with Minimal Semantic Distortion for IP Protection
von: Dang, Kieu, et al.
Veröffentlicht: (2026)
von: Dang, Kieu, et al.
Veröffentlicht: (2026)
OneShield -- the Next Generation of LLM Guardrails
von: DeLuca, Chad, et al.
Veröffentlicht: (2025)
von: DeLuca, Chad, et al.
Veröffentlicht: (2025)
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework
von: Liang, Zi, et al.
Veröffentlicht: (2025)
von: Liang, Zi, et al.
Veröffentlicht: (2025)
Breaking Down the Defenses: A Comparative Survey of Attacks on Large Language Models
von: Chowdhury, Arijit Ghosh, et al.
Veröffentlicht: (2024)
von: Chowdhury, Arijit Ghosh, et al.
Veröffentlicht: (2024)
Improving LLM Safety Alignment with Dual-Objective Optimization
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
Analysing the Safety Pitfalls of Steering Vectors
von: Li, Yuxiao, et al.
Veröffentlicht: (2026)
von: Li, Yuxiao, et al.
Veröffentlicht: (2026)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
Overriding Safety protections of Open-source Models
von: Kumar, Sachin
Veröffentlicht: (2024)
von: Kumar, Sachin
Veröffentlicht: (2024)
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
von: Nawal, Aditya, et al.
Veröffentlicht: (2026)
von: Nawal, Aditya, et al.
Veröffentlicht: (2026)
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
von: Chen, Shuo, et al.
Veröffentlicht: (2025)
von: Chen, Shuo, et al.
Veröffentlicht: (2025)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
von: Fu, Yu, et al.
Veröffentlicht: (2024)
von: Fu, Yu, et al.
Veröffentlicht: (2024)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
von: He, Luxi, et al.
Veröffentlicht: (2024)
von: He, Luxi, et al.
Veröffentlicht: (2024)
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025)
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025)
S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
von: Yuan, Xiaohan, et al.
Veröffentlicht: (2024)
von: Yuan, Xiaohan, et al.
Veröffentlicht: (2024)
LLM Reinforcement in Context
von: Rivasseau, Thomas
Veröffentlicht: (2025)
von: Rivasseau, Thomas
Veröffentlicht: (2025)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
Watermarking LLM Agent Trajectories
von: Meng, Wenlong, et al.
Veröffentlicht: (2026)
von: Meng, Wenlong, et al.
Veröffentlicht: (2026)
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
von: Xie, Yueqi, et al.
Veröffentlicht: (2024)
von: Xie, Yueqi, et al.
Veröffentlicht: (2024)
EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models
von: Wu, Jialin, et al.
Veröffentlicht: (2025)
von: Wu, Jialin, et al.
Veröffentlicht: (2025)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
von: Huang, Caishuang, et al.
Veröffentlicht: (2024)
von: Huang, Caishuang, et al.
Veröffentlicht: (2024)
One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
von: Li, Linbao, et al.
Veröffentlicht: (2025)
von: Li, Linbao, et al.
Veröffentlicht: (2025)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
von: Chae, Kyubyung, et al.
Veröffentlicht: (2025)
von: Chae, Kyubyung, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models
von: Arif, Samee, et al.
Veröffentlicht: (2026) -
Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
von: Deng, Naihao, et al.
Veröffentlicht: (2025) -
Rethinking Table Instruction Tuning
von: Deng, Naihao, et al.
Veröffentlicht: (2025) -
One Trigger Token Is Enough: A Defense Strategy for Balancing Safety and Usability in Large Language Models
von: Gu, Haoran, et al.
Veröffentlicht: (2025) -
Security Attacks on LLM-based Code Completion Tools
von: Cheng, Wen, et al.
Veröffentlicht: (2024)