Your Text Encoder Can Be An Object-Level Watermarking Controller

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Devulapally, Naresh Kumar, Huang, Mingzhen, Asnani, Vishal, Agarwal, Shruti, Lyu, Siwei, Lokhande, Vishnu Suresh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912278341746688
author Devulapally, Naresh Kumar
Huang, Mingzhen
Asnani, Vishal
Agarwal, Shruti
Lyu, Siwei
Lokhande, Vishnu Suresh
author_facet Devulapally, Naresh Kumar
Huang, Mingzhen
Asnani, Vishal
Agarwal, Shruti
Lyu, Siwei
Lokhande, Vishnu Suresh
contents Invisible watermarking of AI-generated images can help with copyright protection, enabling detection and identification of AI-generated media. In this work, we present a novel approach to watermark images of T2I Latent Diffusion Models (LDMs). By only fine-tuning text token embeddings $W_*$, we enable watermarking in selected objects or parts of the image, offering greater flexibility compared to traditional full-image watermarking. Our method leverages the text encoder's compatibility across various LDMs, allowing plug-and-play integration for different LDMs. Moreover, introducing the watermark early in the encoding stage improves robustness to adversarial perturbations in later stages of the pipeline. Our approach achieves $99\%$ bit accuracy ($48$ bits) with a $10^5 \times$ reduction in model parameters, enabling efficient watermarking.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11945
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Your Text Encoder Can Be An Object-Level Watermarking Controller
Devulapally, Naresh Kumar
Huang, Mingzhen
Asnani, Vishal
Agarwal, Shruti
Lyu, Siwei
Lokhande, Vishnu Suresh
Computer Vision and Pattern Recognition
Cryptography and Security
Machine Learning
Invisible watermarking of AI-generated images can help with copyright protection, enabling detection and identification of AI-generated media. In this work, we present a novel approach to watermark images of T2I Latent Diffusion Models (LDMs). By only fine-tuning text token embeddings $W_*$, we enable watermarking in selected objects or parts of the image, offering greater flexibility compared to traditional full-image watermarking. Our method leverages the text encoder's compatibility across various LDMs, allowing plug-and-play integration for different LDMs. Moreover, introducing the watermark early in the encoding stage improves robustness to adversarial perturbations in later stages of the pipeline. Our approach achieves $99\%$ bit accuracy ($48$ bits) with a $10^5 \times$ reduction in model parameters, enabling efficient watermarking.
title Your Text Encoder Can Be An Object-Level Watermarking Controller
topic Computer Vision and Pattern Recognition
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2503.11945