Chain-of-Trigger: An Agentic Backdoor that Paradoxically Enhances Agentic Robustness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiu, Jiyang, Ma, Xinbei, Xu, Yunqing, Zhang, Zhuosheng, Zhao, Hai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916999363297280
author Qiu, Jiyang
Ma, Xinbei
Xu, Yunqing
Zhang, Zhuosheng
Zhao, Hai
author_facet Qiu, Jiyang
Ma, Xinbei
Xu, Yunqing
Zhang, Zhuosheng
Zhao, Hai
contents The rapid deployment of large language model (LLM)-based agents in real-world applications has raised serious concerns about their trustworthiness. In this work, we reveal the security and robustness vulnerabilities of these agents through backdoor attacks. Distinct from traditional backdoors limited to single-step control, we propose the Chain-of-Trigger Backdoor (CoTri), a multi-step backdoor attack designed for long-horizon agentic control. CoTri relies on an ordered sequence. It starts with an initial trigger, and subsequent ones are drawn from the environment, allowing multi-step manipulation that diverts the agent from its intended task. Experimental results show that CoTri achieves a near-perfect attack success rate (ASR) while maintaining a near-zero false trigger rate (FTR). Due to training data modeling the stochastic nature of the environment, the implantation of CoTri paradoxically enhances the agent's performance on benign tasks and even improves its robustness against environmental distractions. We further validate CoTri on vision-language models (VLMs), confirming its scalability to multimodal agents. Our work highlights that CoTri achieves stable, multi-step control within agents, improving their inherent robustness and task capabilities, which ultimately makes the attack more stealthy and raises potential safty risks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08238
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chain-of-Trigger: An Agentic Backdoor that Paradoxically Enhances Agentic Robustness
Qiu, Jiyang
Ma, Xinbei
Xu, Yunqing
Zhang, Zhuosheng
Zhao, Hai
Artificial Intelligence
The rapid deployment of large language model (LLM)-based agents in real-world applications has raised serious concerns about their trustworthiness. In this work, we reveal the security and robustness vulnerabilities of these agents through backdoor attacks. Distinct from traditional backdoors limited to single-step control, we propose the Chain-of-Trigger Backdoor (CoTri), a multi-step backdoor attack designed for long-horizon agentic control. CoTri relies on an ordered sequence. It starts with an initial trigger, and subsequent ones are drawn from the environment, allowing multi-step manipulation that diverts the agent from its intended task. Experimental results show that CoTri achieves a near-perfect attack success rate (ASR) while maintaining a near-zero false trigger rate (FTR). Due to training data modeling the stochastic nature of the environment, the implantation of CoTri paradoxically enhances the agent's performance on benign tasks and even improves its robustness against environmental distractions. We further validate CoTri on vision-language models (VLMs), confirming its scalability to multimodal agents. Our work highlights that CoTri achieves stable, multi-step control within agents, improving their inherent robustness and task capabilities, which ultimately makes the attack more stealthy and raises potential safty risks.
title Chain-of-Trigger: An Agentic Backdoor that Paradoxically Enhances Agentic Robustness
topic Artificial Intelligence
url https://arxiv.org/abs/2510.08238