From Defender to Devil? Unintended Risk Interactions Induced by LLM Defenses
Fuente:
arXiv
Saved in:
| Main Authors: | Meng, Xiangtao, Cong, Tianshuo, Wang, Li, Chen, Wenyu, Li, Zheng, Guo, Shanqing, Wang, Xiaoyun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models
by: Meng, Xiangtao, et al.
Published: (2026)
by: Meng, Xiangtao, et al.
Published: (2026)
Beyond Known Fakes: Generalized Detection of AI-Generated Images via Post-hoc Distribution Alignment
by: Wang, Li, et al.
Published: (2025)
by: Wang, Li, et al.
Published: (2025)
Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models
by: Dong, Yingkai, et al.
Published: (2024)
by: Dong, Yingkai, et al.
Published: (2024)
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
by: Chen, Wenyu, et al.
Published: (2026)
by: Chen, Wenyu, et al.
Published: (2026)
DCMI: A Differential Calibration Membership Inference Attack Against Retrieval-Augmented Generation
by: Gao, Xinyu, et al.
Published: (2025)
by: Gao, Xinyu, et al.
Published: (2025)
SoK: Unintended Interactions among Machine Learning Defenses and Risks
by: Duddu, Vasisht, et al.
Published: (2023)
by: Duddu, Vasisht, et al.
Published: (2023)
AVA: Inconspicuous Attribute Variation-based Adversarial Attack bypassing DeepFake Detection
by: Meng, Xiangtao, et al.
Published: (2023)
by: Meng, Xiangtao, et al.
Published: (2023)
Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification
by: Meng, Xiangtao, et al.
Published: (2025)
by: Meng, Xiangtao, et al.
Published: (2025)
VidLeaks: Membership Inference Attacks Against Text-to-Video Models
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
Cryptanalysis of Pseudorandom Error-Correcting Codes
by: Wang, Tianrui, et al.
Published: (2025)
by: Wang, Tianrui, et al.
Published: (2025)
Watermarking LLM-Generated Datasets in Downstream Tasks
by: Liu, Yugeng, et al.
Published: (2025)
by: Liu, Yugeng, et al.
Published: (2025)
Defending Against Prompt Injection With a Few DefensiveTokens
by: Chen, Sizhe, et al.
Published: (2025)
by: Chen, Sizhe, et al.
Published: (2025)
ICL-EVADER: Zero-Query Black-Box Evasion Attacks on In-Context Learning and Their Defenses
by: He, Ningyuan, et al.
Published: (2026)
by: He, Ningyuan, et al.
Published: (2026)
LoRA-Leak: Membership Inference Attacks Against LoRA Fine-tuned Language Models
by: Ran, Delong, et al.
Published: (2025)
by: Ran, Delong, et al.
Published: (2025)
A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures
by: Kong, Dezhang, et al.
Published: (2025)
by: Kong, Dezhang, et al.
Published: (2025)
United We Defend: Collaborative Membership Inference Defenses in Federated Learning
by: Bai, Li, et al.
Published: (2026)
by: Bai, Li, et al.
Published: (2026)
Conditional Cube Attack on Round-Reduced ASCON
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation
by: Zhang, Shenyi, et al.
Published: (2025)
by: Zhang, Shenyi, et al.
Published: (2025)
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
by: Yan, Lecheng, et al.
Published: (2026)
by: Yan, Lecheng, et al.
Published: (2026)
AGNNCert: Defending Graph Neural Networks against Arbitrary Perturbations with Deterministic Certification
by: Li, Jiate, et al.
Published: (2025)
by: Li, Jiate, et al.
Published: (2025)
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
by: Pang, Yan, et al.
Published: (2025)
by: Pang, Yan, et al.
Published: (2025)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
by: Zheng, Jingyi, et al.
Published: (2024)
by: Zheng, Jingyi, et al.
Published: (2024)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025)
by: Zhang, Heyi, et al.
Published: (2025)
COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers
by: Wang, Junyu, et al.
Published: (2025)
by: Wang, Junyu, et al.
Published: (2025)
FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
by: Gong, Yichen, et al.
Published: (2023)
by: Gong, Yichen, et al.
Published: (2023)
The Devil Behind the Mirror: Tracking the Campaigns of Cryptocurrency Abuses on the Dark Web
by: Xia, Pengcheng, et al.
Published: (2024)
by: Xia, Pengcheng, et al.
Published: (2024)
Poisoning Attacks to Local Differential Privacy for Ranking Estimation
by: Zhan, Pei, et al.
Published: (2025)
by: Zhan, Pei, et al.
Published: (2025)
Defending Against Prompt Injection with DataFilter
by: Wang, Yizhu, et al.
Published: (2025)
by: Wang, Yizhu, et al.
Published: (2025)
Have You Merged My Model? On The Robustness of Large Language Model IP Protection Methods Against Model Merging
by: Cong, Tianshuo, et al.
Published: (2024)
by: Cong, Tianshuo, et al.
Published: (2024)
Dissecting Open Edge Computing Platforms: Ecosystem, Usage, and Security Risks
by: Bi, Yu, et al.
Published: (2024)
by: Bi, Yu, et al.
Published: (2024)
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
by: Wang, Xunguang, et al.
Published: (2024)
by: Wang, Xunguang, et al.
Published: (2024)
Proactive Hardening of LLM Defenses with HASTE
by: Chen, Henry, et al.
Published: (2026)
by: Chen, Henry, et al.
Published: (2026)
SUAD: Solid-Channel Ultrasound Injection Attack and Defense to Voice Assistants
by: Liu, Chao, et al.
Published: (2025)
by: Liu, Chao, et al.
Published: (2025)
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
by: Ying, Zonghao, et al.
Published: (2026)
by: Ying, Zonghao, et al.
Published: (2026)
No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks
by: Leong, Chak Tou, et al.
Published: (2024)
by: Leong, Chak Tou, et al.
Published: (2024)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
by: Campbell, David, et al.
Published: (2026)
by: Campbell, David, et al.
Published: (2026)
Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
by: Zhang, Yixiang, et al.
Published: (2025)
by: Zhang, Yixiang, et al.
Published: (2025)
KinGuard: Hierarchical Kinship-Aware Fingerprinting to Defend Against Large Language Model Stealing
by: Xu, Zhenhua, et al.
Published: (2026)
by: Xu, Zhenhua, et al.
Published: (2026)
Following Devils' Footprint: Towards Real-time Detection of Price Manipulation Attacks
by: Zhang, Bosi, et al.
Published: (2025)
by: Zhang, Bosi, et al.
Published: (2025)
Reframing LLM Agent Security as an Agent-Human Interaction Problem
by: Wang, Peiran, et al.
Published: (2026)
by: Wang, Peiran, et al.
Published: (2026)
Similar Items
-
Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models
by: Meng, Xiangtao, et al.
Published: (2026) -
Beyond Known Fakes: Generalized Detection of AI-Generated Images via Post-hoc Distribution Alignment
by: Wang, Li, et al.
Published: (2025) -
Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models
by: Dong, Yingkai, et al.
Published: (2024) -
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
by: Chen, Wenyu, et al.
Published: (2026) -
DCMI: A Differential Calibration Membership Inference Attack Against Retrieval-Augmented Generation
by: Gao, Xinyu, et al.
Published: (2025)