IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Liwen, Wang, Wenxuan, Wang, Shuai, Li, Zongjie, Ji, Zhenlan, Lyu, Zongyi, Wu, Daoyuan, Cheung, Shing-Chi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908410393395200
author Wang, Liwen
Wang, Wenxuan
Wang, Shuai
Li, Zongjie
Ji, Zhenlan
Lyu, Zongyi
Wu, Daoyuan
Cheung, Shing-Chi
author_facet Wang, Liwen
Wang, Wenxuan
Wang, Shuai
Li, Zongjie
Ji, Zhenlan
Lyu, Zongyi
Wu, Daoyuan
Cheung, Shing-Chi
contents The rapid advancement of Large Language Models (LLMs) has led to the emergence of Multi-Agent Systems (MAS) to perform complex tasks through collaboration. However, the intricate nature of MAS, including their architecture and agent interactions, raises significant concerns regarding intellectual property (IP) protection. In this paper, we introduce MASLEAK, a novel attack framework designed to extract sensitive information from MAS applications. MASLEAK targets a practical, black-box setting, where the adversary has no prior knowledge of the MAS architecture or agent configurations. The adversary can only interact with the MAS through its public API, submitting attack query $q$ and observing outputs from the final agent. Inspired by how computer worms propagate and infect vulnerable network hosts, MASLEAK carefully crafts adversarial query $q$ to elicit, propagate, and retain responses from each MAS agent that reveal a full set of proprietary components, including the number of agents, system topology, system prompts, task instructions, and tool usages. We construct the first synthetic dataset of MAS applications with 810 applications and also evaluate MASLEAK against real-world MAS applications, including Coze and CrewAI. MASLEAK achieves high accuracy in extracting MAS IP, with an average attack success rate of 87% for system prompts and task instructions, and 92% for system architecture in most cases. We conclude by discussing the implications of our findings and the potential defenses.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12442
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
Wang, Liwen
Wang, Wenxuan
Wang, Shuai
Li, Zongjie
Ji, Zhenlan
Lyu, Zongyi
Wu, Daoyuan
Cheung, Shing-Chi
Cryptography and Security
Artificial Intelligence
Computation and Language
The rapid advancement of Large Language Models (LLMs) has led to the emergence of Multi-Agent Systems (MAS) to perform complex tasks through collaboration. However, the intricate nature of MAS, including their architecture and agent interactions, raises significant concerns regarding intellectual property (IP) protection. In this paper, we introduce MASLEAK, a novel attack framework designed to extract sensitive information from MAS applications. MASLEAK targets a practical, black-box setting, where the adversary has no prior knowledge of the MAS architecture or agent configurations. The adversary can only interact with the MAS through its public API, submitting attack query $q$ and observing outputs from the final agent. Inspired by how computer worms propagate and infect vulnerable network hosts, MASLEAK carefully crafts adversarial query $q$ to elicit, propagate, and retain responses from each MAS agent that reveal a full set of proprietary components, including the number of agents, system topology, system prompts, task instructions, and tool usages. We construct the first synthetic dataset of MAS applications with 810 applications and also evaluate MASLEAK against real-world MAS applications, including Coze and CrewAI. MASLEAK achieves high accuracy in extracting MAS IP, with an average attack success rate of 87% for system prompts and task instructions, and 92% for system architecture in most cases. We conclude by discussing the implications of our findings and the potential defenses.
title IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.12442