Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Chenxi, Huang, Ruiyang, Sun, Jiayan, Wei, Lei, Wu, Yifan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917538838872064
author Wang, Chenxi
Huang, Ruiyang
Sun, Jiayan
Wei, Lei
Wu, Yifan
author_facet Wang, Chenxi
Huang, Ruiyang
Sun, Jiayan
Wei, Lei
Wu, Yifan
contents Latent-based multi-agent systems replace parts of explicit inter-agent communication with hidden representations, offering a new direction for efficient and flexible agent collaboration. However, moving coordination into latent space may also move attacks beyond the reach of visible-text inspection. In this paper, we study whether latent states can carry attack-associated information that remains effective during clean executions. To examine this question, we introduce a latent attack framework that reactivates attack-induced effects through latent interventions without reusing adversarial text. Extensive experiments show that the resulting latent-only attacks can substantially degrade task performance in clean executions, especially when applied to inter-agent KV-cache handoffs rather than local hidden states. Further control analyses indicate that this degradation cannot be reduced to arbitrary perturbations or invalid generation. Overall, our findings suggest that latent-based collaboration does not remove attack risk. It shifts part of the risk into less observable execution states, calling for safeguards beyond visible-text inspection.
format Preprint
id arxiv_https___arxiv_org_abs_2605_28214
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems
Wang, Chenxi
Huang, Ruiyang
Sun, Jiayan
Wei, Lei
Wu, Yifan
Cryptography and Security
Machine Learning
Multiagent Systems
Latent-based multi-agent systems replace parts of explicit inter-agent communication with hidden representations, offering a new direction for efficient and flexible agent collaboration. However, moving coordination into latent space may also move attacks beyond the reach of visible-text inspection. In this paper, we study whether latent states can carry attack-associated information that remains effective during clean executions. To examine this question, we introduce a latent attack framework that reactivates attack-induced effects through latent interventions without reusing adversarial text. Extensive experiments show that the resulting latent-only attacks can substantially degrade task performance in clean executions, especially when applied to inter-agent KV-cache handoffs rather than local hidden states. Further control analyses indicate that this degradation cannot be reduced to arbitrary perturbations or invalid generation. Overall, our findings suggest that latent-based collaboration does not remove attack risk. It shifts part of the risk into less observable execution states, calling for safeguards beyond visible-text inspection.
title Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems
topic Cryptography and Security
Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2605.28214