PEAR: Planner-Executor Agent Robustness Benchmark

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dong, Shen, Zhang, Mingxuan, He, Pengfei, Ma, Li, Thuraisingham, Bhavani, Liu, Hui, Xing, Yue
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909984052215808
author Dong, Shen
Zhang, Mingxuan
He, Pengfei
Ma, Li
Thuraisingham, Bhavani
Liu, Hui
Xing, Yue
author_facet Dong, Shen
Zhang, Mingxuan
He, Pengfei
Ma, Li
Thuraisingham, Bhavani
Liu, Hui
Xing, Yue
contents Large Language Model (LLM)-based Multi-Agent Systems (MAS) have emerged as a powerful paradigm for tackling complex, multi-step tasks across diverse domains. However, despite their impressive capabilities, MAS remain susceptible to adversarial manipulation. Existing studies typically examine isolated attack surfaces or specific scenarios, leaving a lack of holistic understanding of MAS vulnerabilities. To bridge this gap, we introduce PEAR, a benchmark for systematically evaluating both the utility and vulnerability of planner-executor MAS. While compatible with various MAS architectures, our benchmark focuses on the planner-executor structure, which is a practical and widely adopted design. Through extensive experiments, we find that (1) a weak planner degrades overall clean task performance more severely than a weak executor; (2) while a memory module is essential for the planner, having a memory module for the executor does not impact the clean task performance; (3) there exists a trade-off between task performance and robustness; and (4) attacks targeting the planner are particularly effective at misleading the system. These findings offer actionable insights for enhancing the robustness of MAS and lay the groundwork for principled defenses in multi-agent settings.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07505
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PEAR: Planner-Executor Agent Robustness Benchmark
Dong, Shen
Zhang, Mingxuan
He, Pengfei
Ma, Li
Thuraisingham, Bhavani
Liu, Hui
Xing, Yue
Machine Learning
Large Language Model (LLM)-based Multi-Agent Systems (MAS) have emerged as a powerful paradigm for tackling complex, multi-step tasks across diverse domains. However, despite their impressive capabilities, MAS remain susceptible to adversarial manipulation. Existing studies typically examine isolated attack surfaces or specific scenarios, leaving a lack of holistic understanding of MAS vulnerabilities. To bridge this gap, we introduce PEAR, a benchmark for systematically evaluating both the utility and vulnerability of planner-executor MAS. While compatible with various MAS architectures, our benchmark focuses on the planner-executor structure, which is a practical and widely adopted design. Through extensive experiments, we find that (1) a weak planner degrades overall clean task performance more severely than a weak executor; (2) while a memory module is essential for the planner, having a memory module for the executor does not impact the clean task performance; (3) there exists a trade-off between task performance and robustness; and (4) attacks targeting the planner are particularly effective at misleading the system. These findings offer actionable insights for enhancing the robustness of MAS and lay the groundwork for principled defenses in multi-agent settings.
title PEAR: Planner-Executor Agent Robustness Benchmark
topic Machine Learning
url https://arxiv.org/abs/2510.07505