Adversarial Attack on Black-Box Multi-Agent by Adaptive Perturbation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Jianming, Wang, Yawen, Wang, Junjie, Xie, Xiaofei, Hu, Yuanzhe, Wang, Qing, Xu, Fanjiang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911627460214784
author Chen, Jianming
Wang, Yawen
Wang, Junjie
Xie, Xiaofei
Hu, Yuanzhe
Wang, Qing
Xu, Fanjiang
author_facet Chen, Jianming
Wang, Yawen
Wang, Junjie
Xie, Xiaofei
Hu, Yuanzhe
Wang, Qing
Xu, Fanjiang
contents Evaluating security and reliability for multi-agent systems (MAS) is urgent as they become increasingly prevalent in various applications. As an evaluation technique, existing adversarial attack frameworks face certain limitations, e.g., impracticality due to the requirement of white-box information or high control authority, and a lack of stealthiness or effectiveness as they often target all agents or specific fixed agents. To address these issues, we propose AdapAM, a novel framework for adversarial attacks on black-box MAS. AdapAM incorporates two key components: (1) Adaptive Selection Policy simultaneously selects the victim and determines the anticipated malicious action (the action would lead to the worst impact on MAS), balancing effectiveness and stealthiness. (2) Proxy-based Perturbation to Induce Malicious Action utilizes generative adversarial imitation learning to approximate the target MAS, allowing AdapAM to generate perturbed observations using white-box information and thus induce victims to execute malicious action in black-box settings. We evaluate AdapAM across eight multi-agent environments and compare it with four state-of-the-art and commonly-used baselines. Results demonstrate that AdapAM achieves the best attack performance in different perturbation rates. Besides, AdapAM-generated perturbations are the least noisy and hardest to detect, emphasizing the stealthiness.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15292
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adversarial Attack on Black-Box Multi-Agent by Adaptive Perturbation
Chen, Jianming
Wang, Yawen
Wang, Junjie
Xie, Xiaofei
Hu, Yuanzhe
Wang, Qing
Xu, Fanjiang
Multiagent Systems
Evaluating security and reliability for multi-agent systems (MAS) is urgent as they become increasingly prevalent in various applications. As an evaluation technique, existing adversarial attack frameworks face certain limitations, e.g., impracticality due to the requirement of white-box information or high control authority, and a lack of stealthiness or effectiveness as they often target all agents or specific fixed agents. To address these issues, we propose AdapAM, a novel framework for adversarial attacks on black-box MAS. AdapAM incorporates two key components: (1) Adaptive Selection Policy simultaneously selects the victim and determines the anticipated malicious action (the action would lead to the worst impact on MAS), balancing effectiveness and stealthiness. (2) Proxy-based Perturbation to Induce Malicious Action utilizes generative adversarial imitation learning to approximate the target MAS, allowing AdapAM to generate perturbed observations using white-box information and thus induce victims to execute malicious action in black-box settings. We evaluate AdapAM across eight multi-agent environments and compare it with four state-of-the-art and commonly-used baselines. Results demonstrate that AdapAM achieves the best attack performance in different perturbation rates. Besides, AdapAM-generated perturbations are the least noisy and hardest to detect, emphasizing the stealthiness.
title Adversarial Attack on Black-Box Multi-Agent by Adaptive Perturbation
topic Multiagent Systems
url https://arxiv.org/abs/2511.15292