Leveraging Optimization for Adaptive Attacks on Image Watermarks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lukas, Nils, Diaa, Abdulrahman, Fenaux, Lucas, Kerschbaum, Florian
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909077969305600
author Lukas, Nils
Diaa, Abdulrahman
Fenaux, Lucas
Kerschbaum, Florian
author_facet Lukas, Nils
Diaa, Abdulrahman
Fenaux, Lucas
Kerschbaum, Florian
contents Untrustworthy users can misuse image generators to synthesize high-quality deepfakes and engage in unethical activities. Watermarking deters misuse by marking generated content with a hidden message, enabling its detection using a secret watermarking key. A core security property of watermarking is robustness, which states that an attacker can only evade detection by substantially degrading image quality. Assessing robustness requires designing an adaptive attack for the specific watermarking algorithm. When evaluating watermarking algorithms and their (adaptive) attacks, it is challenging to determine whether an adaptive attack is optimal, i.e., the best possible attack. We solve this problem by defining an objective function and then approach adaptive attacks as an optimization problem. The core idea of our adaptive attacks is to replicate secret watermarking keys locally by creating surrogate keys that are differentiable and can be used to optimize the attack's parameters. We demonstrate for Stable Diffusion models that such an attacker can break all five surveyed watermarking methods at no visible degradation in image quality. Optimizing our attacks is efficient and requires less than 1 GPU hour to reduce the detection accuracy to 6.3% or less. Our findings emphasize the need for more rigorous robustness testing against adaptive, learnable attackers.
format Preprint
id arxiv_https___arxiv_org_abs_2309_16952
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Leveraging Optimization for Adaptive Attacks on Image Watermarks
Lukas, Nils
Diaa, Abdulrahman
Fenaux, Lucas
Kerschbaum, Florian
Cryptography and Security
Machine Learning
Untrustworthy users can misuse image generators to synthesize high-quality deepfakes and engage in unethical activities. Watermarking deters misuse by marking generated content with a hidden message, enabling its detection using a secret watermarking key. A core security property of watermarking is robustness, which states that an attacker can only evade detection by substantially degrading image quality. Assessing robustness requires designing an adaptive attack for the specific watermarking algorithm. When evaluating watermarking algorithms and their (adaptive) attacks, it is challenging to determine whether an adaptive attack is optimal, i.e., the best possible attack. We solve this problem by defining an objective function and then approach adaptive attacks as an optimization problem. The core idea of our adaptive attacks is to replicate secret watermarking keys locally by creating surrogate keys that are differentiable and can be used to optimize the attack's parameters. We demonstrate for Stable Diffusion models that such an attacker can break all five surveyed watermarking methods at no visible degradation in image quality. Optimizing our attacks is efficient and requires less than 1 GPU hour to reduce the detection accuracy to 6.3% or less. Our findings emphasize the need for more rigorous robustness testing against adaptive, learnable attackers.
title Leveraging Optimization for Adaptive Attacks on Image Watermarks
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2309.16952