Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Boisvert, Léo, Puri, Abhay, Evuru, Chandra Kiran Reddy, Sepahvand, Nazanin, Chapados, Nicolas, Cappart, Quentin, Lacoste, Alexandre, Dvijotham, Krishnamurthy Dj, Drouin, Alexandre
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913028240310272
author Boisvert, Léo
Puri, Abhay
Evuru, Chandra Kiran Reddy
Sepahvand, Nazanin
Chapados, Nicolas
Cappart, Quentin
Lacoste, Alexandre
Dvijotham, Krishnamurthy Dj
Drouin, Alexandre
author_facet Boisvert, Léo
Puri, Abhay
Evuru, Chandra Kiran Reddy
Sepahvand, Nazanin
Chapados, Nicolas
Cappart, Quentin
Lacoste, Alexandre
Dvijotham, Krishnamurthy Dj
Drouin, Alexandre
contents While finetuning AI agents on interaction data -- such as web browsing or tool use -- improves their capabilities, it also introduces critical security vulnerabilities within the agentic AI supply chain. We show that adversaries can effectively poison the data collection pipeline at multiple stages to embed hard-to-detect backdoors that, when triggered, cause unsafe or malicious behavior. We formalize three realistic threat models across distinct layers of the supply chain: direct poisoning of finetuning data, pre-backdoored base models, and environment poisoning, a novel attack vector that exploits vulnerabilities specific to agentic training pipelines. Evaluated on two widely adopted agentic benchmarks, all three threat models prove effective: poisoning only a small number of demonstrations is sufficient to embed a backdoor that causes an agent to leak confidential user information with over 80\% success.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05159
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
Boisvert, Léo
Puri, Abhay
Evuru, Chandra Kiran Reddy
Sepahvand, Nazanin
Chapados, Nicolas
Cappart, Quentin
Lacoste, Alexandre
Dvijotham, Krishnamurthy Dj
Drouin, Alexandre
Cryptography and Security
Artificial Intelligence
Machine Learning
I.2
While finetuning AI agents on interaction data -- such as web browsing or tool use -- improves their capabilities, it also introduces critical security vulnerabilities within the agentic AI supply chain. We show that adversaries can effectively poison the data collection pipeline at multiple stages to embed hard-to-detect backdoors that, when triggered, cause unsafe or malicious behavior. We formalize three realistic threat models across distinct layers of the supply chain: direct poisoning of finetuning data, pre-backdoored base models, and environment poisoning, a novel attack vector that exploits vulnerabilities specific to agentic training pipelines. Evaluated on two widely adopted agentic benchmarks, all three threat models prove effective: poisoning only a small number of demonstrations is sufficient to embed a backdoor that causes an agent to leak confidential user information with over 80\% success.
title Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
topic Cryptography and Security
Artificial Intelligence
Machine Learning
I.2
url https://arxiv.org/abs/2510.05159