Your Agent Can Defend Itself against Backdoor Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Changjiang, Li, Jiacheng, Liang, Bochuan, Cao, Jinghui, Chen, Ting, Wang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
by: Cao, Bochuan, et al.
Published: (2023)
by: Cao, Bochuan, et al.
Published: (2023)
Defending the Edge: Representative-Attention Defense against Backdoor Attacks in Federated Learning
by: Obioma, Chibueze Peace, et al.
Published: (2025)
by: Obioma, Chibueze Peace, et al.
Published: (2025)
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
by: Cao, Bochuan, et al.
Published: (2025)
by: Cao, Bochuan, et al.
Published: (2025)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
by: Jiang, Tanqiu, et al.
Published: (2024)
by: Jiang, Tanqiu, et al.
Published: (2024)
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
by: Cao, Yuanpu, et al.
Published: (2023)
by: Cao, Yuanpu, et al.
Published: (2023)
Compromising Embodied Agents with Contextual Backdoor Attacks
by: Liu, Aishan, et al.
Published: (2024)
by: Liu, Aishan, et al.
Published: (2024)
Watch the Watcher! Backdoor Attacks on Security-Enhancing Diffusion Models
by: Li, Changjiang, et al.
Published: (2024)
by: Li, Changjiang, et al.
Published: (2024)
GraphRAG under Fire
by: Liang, Jiacheng, et al.
Published: (2025)
by: Liang, Jiacheng, et al.
Published: (2025)
Model Extraction Attacks Revisited
by: Liang, Jiacheng, et al.
Published: (2023)
by: Liang, Jiacheng, et al.
Published: (2023)
Invariant Aggregator for Defending against Federated Backdoor Attacks
by: Wang, Xiaoyang, et al.
Published: (2022)
by: Wang, Xiaoyang, et al.
Published: (2022)
Defending against Backdoor Attack on Deep Neural Networks
by: Cheng, Hao, et al.
Published: (2020)
by: Cheng, Hao, et al.
Published: (2020)
Client-Side Patching against Backdoor Attacks in Federated Learning
by: Molina-Coronado, Borja
Published: (2024)
by: Molina-Coronado, Borja
Published: (2024)
Heterogeneous Graph Backdoor Attack
by: Chen, Jiawei, et al.
Published: (2025)
by: Chen, Jiawei, et al.
Published: (2025)
Data Free Backdoor Attacks
by: Cao, Bochuan, et al.
Published: (2024)
by: Cao, Bochuan, et al.
Published: (2024)
TrojFM: Resource-efficient Backdoor Attacks against Very Large Foundation Models
by: Nie, Yuzhou., et al.
Published: (2024)
by: Nie, Yuzhou., et al.
Published: (2024)
BLAST: A Stealthy Backdoor Leverage Attack against Cooperative Multi-Agent Deep Reinforcement Learning based Systems
by: Fang, Jing, et al.
Published: (2025)
by: Fang, Jing, et al.
Published: (2025)
A Semantic and Clean-label Backdoor Attack against Graph Convolutional Networks
by: Dai, Jiazhu, et al.
Published: (2025)
by: Dai, Jiazhu, et al.
Published: (2025)
FedDefender: Backdoor Attack Defense in Federated Learning
by: Gill, Waris, et al.
Published: (2023)
by: Gill, Waris, et al.
Published: (2023)
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
by: Ren, Zhiyao, et al.
Published: (2025)
by: Ren, Zhiyao, et al.
Published: (2025)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
Invisible Backdoor Attack Through Singular Value Decomposition
by: Chen, Wenmin, et al.
Published: (2024)
by: Chen, Wenmin, et al.
Published: (2024)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Revisiting Backdoor Attacks on Time Series Classification in the Frequency Domain
by: Huang, Yuanmin, et al.
Published: (2025)
by: Huang, Yuanmin, et al.
Published: (2025)
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
by: Pawlak, Stanisław, et al.
Published: (2025)
by: Pawlak, Stanisław, et al.
Published: (2025)
Trapping Attacker in Dilemma: Examining Internal Correlations and External Influences of Trigger for Defending GNN Backdoors
by: Yang, Fan, et al.
Published: (2026)
by: Yang, Fan, et al.
Published: (2026)
Buffer is All You Need: Defending Federated Learning against Backdoor Attacks under Non-iids via Buffering
by: Lyu, Xingyu, et al.
Published: (2025)
by: Lyu, Xingyu, et al.
Published: (2025)
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Structure-Aware Distributed Backdoor Attacks in Federated Learning
by: Jian, Wang, et al.
Published: (2026)
by: Jian, Wang, et al.
Published: (2026)
LoBAM: LoRA-Based Backdoor Attack on Model Merging
by: Yin, Ming, et al.
Published: (2024)
by: Yin, Ming, et al.
Published: (2024)
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
by: Shin, Jeongjin, et al.
Published: (2024)
by: Shin, Jeongjin, et al.
Published: (2024)
Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack
by: Yue, Murong, et al.
Published: (2025)
by: Yue, Murong, et al.
Published: (2025)
Backdoor Attack on Vertical Federated Graph Neural Network Learning
by: Yang, Jirui, et al.
Published: (2024)
by: Yang, Jirui, et al.
Published: (2024)
Learning to Defend by Attacking (and Vice-Versa): Transfer of Learning in Cybersecurity Games
by: Malloy, Tailia, et al.
Published: (2023)
by: Malloy, Tailia, et al.
Published: (2023)
Defending against Adversarial Malware Attacks on ML-based Android Malware Detection Systems
by: He, Ping, et al.
Published: (2025)
by: He, Ping, et al.
Published: (2025)
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
by: Chen, Shuhao, et al.
Published: (2026)
by: Chen, Shuhao, et al.
Published: (2026)
RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models
by: Liang, Jiacheng, et al.
Published: (2026)
by: Liang, Jiacheng, et al.
Published: (2026)
Backdoor Attack against One-Class Sequential Anomaly Detection Models
by: Cheng, He, et al.
Published: (2024)
by: Cheng, He, et al.
Published: (2024)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
by: Wang, Zhun, et al.
Published: (2026)
by: Wang, Zhun, et al.
Published: (2026)
BACKTIME: Backdoor Attacks on Multivariate Time Series Forecasting
by: Lin, Xiao, et al.
Published: (2024)
by: Lin, Xiao, et al.
Published: (2024)
FLARE: Toward Universal Dataset Purification against Backdoor Attacks
by: Hou, Linshan, et al.
Published: (2024)
by: Hou, Linshan, et al.
Published: (2024)
Similar Items
-
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
by: Cao, Bochuan, et al.
Published: (2023) -
Defending the Edge: Representative-Attention Defense against Backdoor Attacks in Federated Learning
by: Obioma, Chibueze Peace, et al.
Published: (2025) -
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
by: Cao, Bochuan, et al.
Published: (2025) -
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
by: Jiang, Tanqiu, et al.
Published: (2024) -
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
by: Cao, Yuanpu, et al.
Published: (2023)