Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913028240310272 |
|---|---|
| author | Boisvert, Léo Puri, Abhay Evuru, Chandra Kiran Reddy Sepahvand, Nazanin Chapados, Nicolas Cappart, Quentin Lacoste, Alexandre Dvijotham, Krishnamurthy Dj Drouin, Alexandre |
| author_facet | Boisvert, Léo Puri, Abhay Evuru, Chandra Kiran Reddy Sepahvand, Nazanin Chapados, Nicolas Cappart, Quentin Lacoste, Alexandre Dvijotham, Krishnamurthy Dj Drouin, Alexandre |
| contents | While finetuning AI agents on interaction data -- such as web browsing or tool use -- improves their capabilities, it also introduces critical security vulnerabilities within the agentic AI supply chain. We show that adversaries can effectively poison the data collection pipeline at multiple stages to embed hard-to-detect backdoors that, when triggered, cause unsafe or malicious behavior. We formalize three realistic threat models across distinct layers of the supply chain: direct poisoning of finetuning data, pre-backdoored base models, and environment poisoning, a novel attack vector that exploits vulnerabilities specific to agentic training pipelines. Evaluated on two widely adopted agentic benchmarks, all three threat models prove effective: poisoning only a small number of demonstrations is sufficient to embed a backdoor that causes an agent to leak confidential user information with over 80\% success. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_05159 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain Boisvert, Léo Puri, Abhay Evuru, Chandra Kiran Reddy Sepahvand, Nazanin Chapados, Nicolas Cappart, Quentin Lacoste, Alexandre Dvijotham, Krishnamurthy Dj Drouin, Alexandre Cryptography and Security Artificial Intelligence Machine Learning I.2 While finetuning AI agents on interaction data -- such as web browsing or tool use -- improves their capabilities, it also introduces critical security vulnerabilities within the agentic AI supply chain. We show that adversaries can effectively poison the data collection pipeline at multiple stages to embed hard-to-detect backdoors that, when triggered, cause unsafe or malicious behavior. We formalize three realistic threat models across distinct layers of the supply chain: direct poisoning of finetuning data, pre-backdoored base models, and environment poisoning, a novel attack vector that exploits vulnerabilities specific to agentic training pipelines. Evaluated on two widely adopted agentic benchmarks, all three threat models prove effective: poisoning only a small number of demonstrations is sufficient to embed a backdoor that causes an agent to leak confidential user information with over 80\% success. |
| title | Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain |
| topic | Cryptography and Security Artificial Intelligence Machine Learning I.2 |
| url | https://arxiv.org/abs/2510.05159 |