Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Draganov, Andrew, Dur, Tolga H., Bhongade, Anandmayi, Phuong, Mary
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908813996589056
author Draganov, Andrew
Dur, Tolga H.
Bhongade, Anandmayi
Phuong, Mary
author_facet Draganov, Andrew
Dur, Tolga H.
Bhongade, Anandmayi
Phuong, Mary
contents We present a data poisoning attack -- Phantom Transfer -- with the property that, even if you know precisely how the poison was placed into an otherwise benign dataset, you cannot filter it out. We achieve this by modifying subliminal learning to work in real-world contexts and demonstrate that the attack works across models, including GPT-4.1. Indeed, even fully paraphrasing every sample in the dataset using a different model does not stop the attack. We also discuss connections to steering vectors and show that one can plant password-triggered behaviours into models while still beating defences. This suggests that data-level defences are insufficient for stopping sophisticated data poisoning attacks. We suggest that future work should focus on model audits and white-box security methods.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04899
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoning
Draganov, Andrew
Dur, Tolga H.
Bhongade, Anandmayi
Phuong, Mary
Cryptography and Security
Artificial Intelligence
We present a data poisoning attack -- Phantom Transfer -- with the property that, even if you know precisely how the poison was placed into an otherwise benign dataset, you cannot filter it out. We achieve this by modifying subliminal learning to work in real-world contexts and demonstrate that the attack works across models, including GPT-4.1. Indeed, even fully paraphrasing every sample in the dataset using a different model does not stop the attack. We also discuss connections to steering vectors and show that one can plant password-triggered behaviours into models while still beating defences. This suggests that data-level defences are insufficient for stopping sophisticated data poisoning attacks. We suggest that future work should focus on model audits and white-box security methods.
title Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoning
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2602.04899