Safeguarding Text-to-Image Generative Models Against Unauthorized Knowledge Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Yilan, Huang, Sida, Zhang, Hongyuan, Li, Xuelong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910245137154048
author Gao, Yilan
Huang, Sida
Zhang, Hongyuan
Li, Xuelong
author_facet Gao, Yilan
Huang, Sida
Zhang, Hongyuan
Li, Xuelong
contents Closed-weight generative services are increasingly deployed through query-based APIs, where users can obtain generated outputs while model parameters remain inaccessible. However, such deployment does not prevent model stealing: an attacker can repeatedly query the service, collect large volumes of released synthetic images, and use them as training data for a private substitute model. This query-output-driven process enables unauthorized knowledge distillation and capability replication without direct access to the original weights. To mitigate this threat, a practical defense should preserve the visual fidelity of released images, provide explicit control over perturbation magnitude, and scale efficiently to large-volume output release. We present WaveGuard, a single-pass, generator-based protection framework that safeguards released synthetic images under a user-specified perturbation budget. WaveGuard employs a frequency-aware perturbation generator to inject structured, imperceptible perturbations that maintain perceptual utility for benign viewers while reducing the usefulness of protected images as training data for unauthorized student models. Extensive experiments under WikiArt-related synthetic-output distillation settings show that WaveGuard achieves a favorable efficacy--fidelity--efficiency trade-off, with explicit imperceptibility control and substantial gains in protection efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22060
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Safeguarding Text-to-Image Generative Models Against Unauthorized Knowledge Distillation
Gao, Yilan
Huang, Sida
Zhang, Hongyuan
Li, Xuelong
Cryptography and Security
Artificial Intelligence
Closed-weight generative services are increasingly deployed through query-based APIs, where users can obtain generated outputs while model parameters remain inaccessible. However, such deployment does not prevent model stealing: an attacker can repeatedly query the service, collect large volumes of released synthetic images, and use them as training data for a private substitute model. This query-output-driven process enables unauthorized knowledge distillation and capability replication without direct access to the original weights. To mitigate this threat, a practical defense should preserve the visual fidelity of released images, provide explicit control over perturbation magnitude, and scale efficiently to large-volume output release. We present WaveGuard, a single-pass, generator-based protection framework that safeguards released synthetic images under a user-specified perturbation budget. WaveGuard employs a frequency-aware perturbation generator to inject structured, imperceptible perturbations that maintain perceptual utility for benign viewers while reducing the usefulness of protected images as training data for unauthorized student models. Extensive experiments under WikiArt-related synthetic-output distillation settings show that WaveGuard achieves a favorable efficacy--fidelity--efficiency trade-off, with explicit imperceptibility control and substantial gains in protection efficiency.
title Safeguarding Text-to-Image Generative Models Against Unauthorized Knowledge Distillation
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2605.22060