AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Tanqiu, Wang, Yuhui, Liang, Jiacheng, Wang, Ting
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911456793985024
author Jiang, Tanqiu
Wang, Yuhui
Liang, Jiacheng
Wang, Ting
author_facet Jiang, Tanqiu
Wang, Yuhui
Liang, Jiacheng
Wang, Ting
contents LLM agents are increasingly deployed in long-horizon, complex environments to solve challenging problems, but this expansion exposes them to long-horizon attacks that exploit multi-turn user-agent-environment interactions to achieve objectives infeasible in single-turn settings. To measure agent vulnerabilities to such risks, we present AgentLAB, the first benchmark dedicated to evaluating LLM agent susceptibility to adaptive, long-horizon attacks. Currently, AgentLAB supports five novel attack types including intent hijacking, tool chaining, task injection, objective drifting, and memory poisoning, spanning 28 realistic agentic environments, and 644 security test cases. Leveraging AgentLAB, we evaluate representative LLM agents and find that they remain highly susceptible to long-horizon attacks; moreover, defenses designed for single-turn interactions fail to reliably mitigate long-horizon threats. We anticipate that AgentLAB will serve as a valuable benchmark for tracking progress on securing LLM agents in practical settings. The benchmark is publicly available at https://tanqiujiang.github.io/AgentLAB_main.
format Preprint
id arxiv_https___arxiv_org_abs_2602_16901
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks
Jiang, Tanqiu
Wang, Yuhui
Liang, Jiacheng
Wang, Ting
Artificial Intelligence
LLM agents are increasingly deployed in long-horizon, complex environments to solve challenging problems, but this expansion exposes them to long-horizon attacks that exploit multi-turn user-agent-environment interactions to achieve objectives infeasible in single-turn settings. To measure agent vulnerabilities to such risks, we present AgentLAB, the first benchmark dedicated to evaluating LLM agent susceptibility to adaptive, long-horizon attacks. Currently, AgentLAB supports five novel attack types including intent hijacking, tool chaining, task injection, objective drifting, and memory poisoning, spanning 28 realistic agentic environments, and 644 security test cases. Leveraging AgentLAB, we evaluate representative LLM agents and find that they remain highly susceptible to long-horizon attacks; moreover, defenses designed for single-turn interactions fail to reliably mitigate long-horizon threats. We anticipate that AgentLAB will serve as a valuable benchmark for tracking progress on securing LLM agents in practical settings. The benchmark is publicly available at https://tanqiujiang.github.io/AgentLAB_main.
title AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks
topic Artificial Intelligence
url https://arxiv.org/abs/2602.16901