The Good, the Bad and the Ugly: Meta-Analysis of Watermarks, Transferable Attacks and Adversarial Defenses

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Głuch, Grzegorz, Turan, Berkant, Nagarajan, Sai Ganesh, Pokutta, Sebastian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917215508365312
author Głuch, Grzegorz
Turan, Berkant
Nagarajan, Sai Ganesh
Pokutta, Sebastian
author_facet Głuch, Grzegorz
Turan, Berkant
Nagarajan, Sai Ganesh
Pokutta, Sebastian
contents We formalize and analyze the trade-off between backdoor-based watermarks and adversarial defenses, framing it as an interactive protocol between a verifier and a prover. While previous works have primarily focused on this trade-off, our analysis extends it by identifying transferable attacks as a third, counterintuitive, but necessary option. Our main result shows that for all learning tasks, at least one of the three exists: a watermark, an adversarial defense, or a transferable attack. By transferable attack, we refer to an efficient algorithm that generates queries indistinguishable from the data distribution and capable of fooling all efficient defenders. Using cryptographic techniques, specifically fully homomorphic encryption, we construct a transferable attack and prove its necessity in this trade-off. Finally, we show that tasks of bounded VC-dimension allow adversarial defenses against all attackers, while a subclass allows watermarks secure against fast adversaries.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08864
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Good, the Bad and the Ugly: Meta-Analysis of Watermarks, Transferable Attacks and Adversarial Defenses
Głuch, Grzegorz
Turan, Berkant
Nagarajan, Sai Ganesh
Pokutta, Sebastian
Machine Learning
Artificial Intelligence
Cryptography and Security
68T01, 94A60, 91A99
We formalize and analyze the trade-off between backdoor-based watermarks and adversarial defenses, framing it as an interactive protocol between a verifier and a prover. While previous works have primarily focused on this trade-off, our analysis extends it by identifying transferable attacks as a third, counterintuitive, but necessary option. Our main result shows that for all learning tasks, at least one of the three exists: a watermark, an adversarial defense, or a transferable attack. By transferable attack, we refer to an efficient algorithm that generates queries indistinguishable from the data distribution and capable of fooling all efficient defenders. Using cryptographic techniques, specifically fully homomorphic encryption, we construct a transferable attack and prove its necessity in this trade-off. Finally, we show that tasks of bounded VC-dimension allow adversarial defenses against all attackers, while a subclass allows watermarks secure against fast adversaries.
title The Good, the Bad and the Ugly: Meta-Analysis of Watermarks, Transferable Attacks and Adversarial Defenses
topic Machine Learning
Artificial Intelligence
Cryptography and Security
68T01, 94A60, 91A99
url https://arxiv.org/abs/2410.08864