LLM Safeguard is a Double-Edged Sword: Exploiting False Positives for Denial-of-Service Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Qingzhao, Xiong, Ziyang, Mao, Z. Morley |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model
by: Wang, Tianyi, et al.
Published: (2026)
by: Wang, Tianyi, et al.
Published: (2026)
From Stealthy Data Fabrication to Unsafe Driving: Realistic Scenario Attacks on Collaborative Perception
by: Zhang, Qingzhao, et al.
Published: (2026)
by: Zhang, Qingzhao, et al.
Published: (2026)
Semantic Denial of Service in LLM-controlled robots
by: Steinberg, Jonathan, et al.
Published: (2026)
by: Steinberg, Jonathan, et al.
Published: (2026)
ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite Thinking
by: Li, Yunzhe, et al.
Published: (2025)
by: Li, Yunzhe, et al.
Published: (2025)
ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models
by: Liu, Xiaogeng, et al.
Published: (2026)
by: Liu, Xiaogeng, et al.
Published: (2026)
CompressionAttack: Exploiting Prompt Compression as a New Attack Surface in LLM-Powered Agents
by: Liu, Zesen, et al.
Published: (2025)
by: Liu, Zesen, et al.
Published: (2025)
Noise as a Double-Edged Sword: Reinforcement Learning Exploits Randomized Defenses in Neural Networks
by: Bakos, Steve, et al.
Published: (2024)
by: Bakos, Steve, et al.
Published: (2024)
DrLLM: Prompt-Enhanced Distributed Denial-of-Service Resistance Method with Large Language Models
by: Yin, Zhenyu, et al.
Published: (2024)
by: Yin, Zhenyu, et al.
Published: (2024)
Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking
by: Li, Yanzeng, et al.
Published: (2025)
by: Li, Yanzeng, et al.
Published: (2025)
Prompt-in-Content Attacks: Exploiting Uploaded Inputs to Hijack LLM Behavior
by: Lian, Zhuotao, et al.
Published: (2025)
by: Lian, Zhuotao, et al.
Published: (2025)
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
by: Li, Siyuan, et al.
Published: (2026)
by: Li, Siyuan, et al.
Published: (2026)
SoK: How Sensor Attacks Disrupt Autonomous Vehicles: An End-to-end Analysis, Challenges, and Missed Threats
by: Zhang, Qingzhao, et al.
Published: (2025)
by: Zhang, Qingzhao, et al.
Published: (2025)
GuardReasoner: Towards Reasoning-based LLM Safeguards
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
What Would Trojans Do? Exploiting Partial-Information Vulnerabilities in Autonomous Vehicle Sensing
by: Hallyburton, R. Spencer, et al.
Published: (2023)
by: Hallyburton, R. Spencer, et al.
Published: (2023)
CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
Prompt-Induced Over-Generation as Denial-of-Service: A Black-Box Attack-Side Benchmark
by: Manu, et al.
Published: (2025)
by: Manu, et al.
Published: (2025)
Causality Laundering: Denial-Feedback Leakage in Tool-Calling LLM Agents
by: Chinaei, Mohammad Hossein
Published: (2026)
by: Chinaei, Mohammad Hossein
Published: (2026)
Safeguarding Multimodal Knowledge Copyright in the RAG-as-a-Service Environment
by: Chen, Tianyu, et al.
Published: (2025)
by: Chen, Tianyu, et al.
Published: (2025)
AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
by: Wu, Yixin, et al.
Published: (2025)
by: Wu, Yixin, et al.
Published: (2025)
Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
by: You, Ziyang, et al.
Published: (2026)
by: You, Ziyang, et al.
Published: (2026)
Exploring and Developing a Pre-Model Safeguard with Draft Models
by: Cai, Hongyu, et al.
Published: (2026)
by: Cai, Hongyu, et al.
Published: (2026)
Breaking the Loop: Detecting and Mitigating Denial-of-Service Vulnerabilities in Large Language Models
by: Yu, Junzhe, et al.
Published: (2025)
by: Yu, Junzhe, et al.
Published: (2025)
Adversarial Reinforcement Learning for Detecting False Data Injection Attacks in Vehicular Routing
by: Eghtesad, Taha, et al.
Published: (2026)
by: Eghtesad, Taha, et al.
Published: (2026)
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
by: Lin, Lixing, et al.
Published: (2026)
by: Lin, Lixing, et al.
Published: (2026)
Depth Gives a False Sense of Privacy: LLM Internal States Inversion
by: Dong, Tian, et al.
Published: (2025)
by: Dong, Tian, et al.
Published: (2025)
Automated and Explainable Denial of Service Analysis for AI-Driven Intrusion Detection Systems
by: Yakubu, Paul Badu, et al.
Published: (2025)
by: Yakubu, Paul Badu, et al.
Published: (2025)
Safeguarding Large Language Models: A Survey
by: Dong, Yi, et al.
Published: (2024)
by: Dong, Yi, et al.
Published: (2024)
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
Enhancing LLM-based Autonomous Driving Agents to Mitigate Perception Attacks
by: Song, Ruoyu, et al.
Published: (2024)
by: Song, Ruoyu, et al.
Published: (2024)
Evaluating False Alarm and Missing Attacks in CAN IDS
by: Hossain, Nirab, et al.
Published: (2026)
by: Hossain, Nirab, et al.
Published: (2026)
Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents
by: Zou, Wei, et al.
Published: (2026)
by: Zou, Wei, et al.
Published: (2026)
LLM Agents can Autonomously Exploit One-day Vulnerabilities
by: Fang, Richard, et al.
Published: (2024)
by: Fang, Richard, et al.
Published: (2024)
PropGuard: Safeguarding LLM-MAS via Propagation-Aware Exploration and Remediation
by: Yan, Bingyu, et al.
Published: (2026)
by: Yan, Bingyu, et al.
Published: (2026)
False Claims against Model Ownership Resolution
by: Liu, Jian, et al.
Published: (2023)
by: Liu, Jian, et al.
Published: (2023)
Compiled Models, Built-In Exploits: Uncovering Pervasive Bit-Flip Attack Surfaces in DNN Executables
by: Chen, Yanzuo, et al.
Published: (2023)
by: Chen, Yanzuo, et al.
Published: (2023)
MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs
by: Ding, Ruyi, et al.
Published: (2025)
by: Ding, Ruyi, et al.
Published: (2025)
Exploiting LLM Quantization
by: Egashira, Kazuki, et al.
Published: (2024)
by: Egashira, Kazuki, et al.
Published: (2024)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
by: Lee, Seunghyun, et al.
Published: (2026)
by: Lee, Seunghyun, et al.
Published: (2026)
On Evaluating the Durability of Safeguards for Open-Weight LLMs
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
PromptKeeper: Safeguarding System Prompts for LLMs
by: Jiang, Zhifeng, et al.
Published: (2024)
by: Jiang, Zhifeng, et al.
Published: (2024)
Similar Items
-
Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model
by: Wang, Tianyi, et al.
Published: (2026) -
From Stealthy Data Fabrication to Unsafe Driving: Realistic Scenario Attacks on Collaborative Perception
by: Zhang, Qingzhao, et al.
Published: (2026) -
Semantic Denial of Service in LLM-controlled robots
by: Steinberg, Jonathan, et al.
Published: (2026) -
ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite Thinking
by: Li, Yunzhe, et al.
Published: (2025) -
ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models
by: Liu, Xiaogeng, et al.
Published: (2026)