BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Yunhao, Li, Yige, Wu, Yutao, Tan, Yingshui, Guo, Yanming, Ding, Yifan, Zhai, Kun, Ma, Xingjun, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
von: Li, Yige, et al.
Veröffentlicht: (2025)
von: Li, Yige, et al.
Veröffentlicht: (2025)
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
von: Li, Juncheng, et al.
Veröffentlicht: (2025)
von: Li, Juncheng, et al.
Veröffentlicht: (2025)
BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
von: Li, Yige, et al.
Veröffentlicht: (2024)
von: Li, Yige, et al.
Veröffentlicht: (2024)
Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks
von: Li, Yige, et al.
Veröffentlicht: (2024)
von: Li, Yige, et al.
Veröffentlicht: (2024)
Expose Before You Defend: Unifying and Enhancing Backdoor Defenses via Exposed Models
von: Li, Yige, et al.
Veröffentlicht: (2024)
von: Li, Yige, et al.
Veröffentlicht: (2024)
BraveGuard: From Open-World Threats to Safer Computer-Use Agents
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
von: Li, Yige, et al.
Veröffentlicht: (2026)
von: Li, Yige, et al.
Veröffentlicht: (2026)
DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agent
von: Zhu, Pengyu, et al.
Veröffentlicht: (2025)
von: Zhu, Pengyu, et al.
Veröffentlicht: (2025)
End-to-End Anti-Backdoor Learning on Images and Time Series
von: Jiang, Yujing, et al.
Veröffentlicht: (2024)
von: Jiang, Yujing, et al.
Veröffentlicht: (2024)
Detecting Backdoor Samples in Contrastive Language Image Pretraining
von: Huang, Hanxun, et al.
Veröffentlicht: (2025)
von: Huang, Hanxun, et al.
Veröffentlicht: (2025)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
von: Jiang, Peihai, et al.
Veröffentlicht: (2025)
von: Jiang, Peihai, et al.
Veröffentlicht: (2025)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
von: Xu, Zonghuan, et al.
Veröffentlicht: (2025)
von: Xu, Zonghuan, et al.
Veröffentlicht: (2025)
HoneypotNet: Backdoor Attacks Against Model Extraction
von: Wang, Yixu, et al.
Veröffentlicht: (2025)
von: Wang, Yixu, et al.
Veröffentlicht: (2025)
Compromising Embodied Agents with Contextual Backdoor Attacks
von: Liu, Aishan, et al.
Veröffentlicht: (2024)
von: Liu, Aishan, et al.
Veröffentlicht: (2024)
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
Stateful Agent Backdoor
von: Dai, Zhengchunmin, et al.
Veröffentlicht: (2026)
von: Dai, Zhengchunmin, et al.
Veröffentlicht: (2026)
Towards Unified Robustness Against Both Backdoor and Adversarial Attacks
von: Niu, Zhenxing, et al.
Veröffentlicht: (2024)
von: Niu, Zhenxing, et al.
Veröffentlicht: (2024)
Collaborative Shadows: Distributed Backdoor Attacks in LLM-Based Multi-Agent Systems
von: Zhu, Pengyu, et al.
Veröffentlicht: (2025)
von: Zhu, Pengyu, et al.
Veröffentlicht: (2025)
Physical Backdoor: Towards Temperature-based Backdoor Attacks in the Physical World
von: Yin, Wen, et al.
Veröffentlicht: (2024)
von: Yin, Wen, et al.
Veröffentlicht: (2024)
Mitigating Backdoor Attack by Injecting Proactive Defensive Backdoor
von: Wei, Shaokui, et al.
Veröffentlicht: (2024)
von: Wei, Shaokui, et al.
Veröffentlicht: (2024)
Backdoor for Debias: Mitigating Model Bias with Backdoor Attack-based Artificial Bias
von: Wu, Shangxi, et al.
Veröffentlicht: (2023)
von: Wu, Shangxi, et al.
Veröffentlicht: (2023)
Your Agent Can Defend Itself against Backdoor Attacks
von: Changjiang, Li, et al.
Veröffentlicht: (2025)
von: Changjiang, Li, et al.
Veröffentlicht: (2025)
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
von: Cui, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Cui, Kaiyuan, et al.
Veröffentlicht: (2026)
AgentRAE: Remote Action Execution through Notification-based Visual Backdoors against Screenshots-based Mobile GUI Agents
von: Luo, Yutao, et al.
Veröffentlicht: (2026)
von: Luo, Yutao, et al.
Veröffentlicht: (2026)
ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
von: Wu, Yutao, et al.
Veröffentlicht: (2025)
VillanDiffusion: A Unified Backdoor Attack Framework for Diffusion Models
von: Chou, Sheng-Yen, et al.
Veröffentlicht: (2023)
von: Chou, Sheng-Yen, et al.
Veröffentlicht: (2023)
Is the Trigger Essential? A Feature-Based Triggerless Backdoor Attack in Vertical Federated Learning
von: Liu, Yige, et al.
Veröffentlicht: (2026)
von: Liu, Yige, et al.
Veröffentlicht: (2026)
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
von: Ge, Huaizhi, et al.
Veröffentlicht: (2024)
von: Ge, Huaizhi, et al.
Veröffentlicht: (2024)
Parasite: A Steganography-based Backdoor Attack Framework for Diffusion Models
von: Chen, Jiahao, et al.
Veröffentlicht: (2025)
von: Chen, Jiahao, et al.
Veröffentlicht: (2025)
Event Trojan: Asynchronous Event-based Backdoor Attacks
von: Wang, Ruofei, et al.
Veröffentlicht: (2024)
von: Wang, Ruofei, et al.
Veröffentlicht: (2024)
Data Free Backdoor Attacks
von: Cao, Bochuan, et al.
Veröffentlicht: (2024)
von: Cao, Bochuan, et al.
Veröffentlicht: (2024)
Can We Trust Embodied Agents? Exploring Backdoor Attacks against Embodied LLM-based Decision-Making Systems
von: Jiao, Ruochen, et al.
Veröffentlicht: (2024)
von: Jiao, Ruochen, et al.
Veröffentlicht: (2024)
Persistent Backdoor Attacks in Continual Learning
von: Guo, Zhen, et al.
Veröffentlicht: (2024)
von: Guo, Zhen, et al.
Veröffentlicht: (2024)
Invisible Backdoor Attacks on Diffusion Models
von: Li, Sen, et al.
Veröffentlicht: (2024)
von: Li, Sen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
von: Feng, Yunhao, et al.
Veröffentlicht: (2026) -
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
von: Feng, Yunhao, et al.
Veröffentlicht: (2026) -
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
von: Li, Yige, et al.
Veröffentlicht: (2025) -
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
von: Li, Juncheng, et al.
Veröffentlicht: (2025) -
BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
von: Li, Yige, et al.
Veröffentlicht: (2024)