Transpose Attack: Stealing Datasets with Bidirectional Training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Amit, Guy, Levy, Mosh, Mirsky, Yisroel
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910450874056704
author Amit, Guy
Levy, Mosh
Mirsky, Yisroel
author_facet Amit, Guy
Levy, Mosh
Mirsky, Yisroel
contents Deep neural networks are normally executed in the forward direction. However, in this work, we identify a vulnerability that enables models to be trained in both directions and on different tasks. Adversaries can exploit this capability to hide rogue models within seemingly legitimate models. In addition, in this work we show that neural networks can be taught to systematically memorize and retrieve specific samples from datasets. Together, these findings expose a novel method in which adversaries can exfiltrate datasets from protected learning environments under the guise of legitimate models. We focus on the data exfiltration attack and show that modern architectures can be used to secretly exfiltrate tens of thousands of samples with high fidelity, high enough to compromise data privacy and even train new models. Moreover, to mitigate this threat we propose a novel approach for detecting infected models.
format Preprint
id arxiv_https___arxiv_org_abs_2311_07389
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Transpose Attack: Stealing Datasets with Bidirectional Training
Amit, Guy
Levy, Mosh
Mirsky, Yisroel
Machine Learning
Cryptography and Security
Deep neural networks are normally executed in the forward direction. However, in this work, we identify a vulnerability that enables models to be trained in both directions and on different tasks. Adversaries can exploit this capability to hide rogue models within seemingly legitimate models. In addition, in this work we show that neural networks can be taught to systematically memorize and retrieve specific samples from datasets. Together, these findings expose a novel method in which adversaries can exfiltrate datasets from protected learning environments under the guise of legitimate models. We focus on the data exfiltration attack and show that modern architectures can be used to secretly exfiltrate tens of thousands of samples with high fidelity, high enough to compromise data privacy and even train new models. Moreover, to mitigate this threat we propose a novel approach for detecting infected models.
title Transpose Attack: Stealing Datasets with Bidirectional Training
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2311.07389