How Many Images Does It Take? Estimating Imitation Thresholds in Text-to-Image Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Verma, Sahil, Rassin, Royi, Das, Arnav, Bhatt, Gantavya, Seshadri, Preethi, Shah, Chirag, Bilmes, Jeff, Hajishirzi, Hannaneh, Elazar, Yanai
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909981819797504
author Verma, Sahil
Rassin, Royi
Das, Arnav
Bhatt, Gantavya
Seshadri, Preethi
Shah, Chirag
Bilmes, Jeff
Hajishirzi, Hannaneh
Elazar, Yanai
author_facet Verma, Sahil
Rassin, Royi
Das, Arnav
Bhatt, Gantavya
Seshadri, Preethi
Shah, Chirag
Bilmes, Jeff
Hajishirzi, Hannaneh
Elazar, Yanai
contents Text-to-image models are trained using large datasets of image-text pairs collected from the internet. These datasets often include copyrighted and private images. Training models on such datasets enables them to generate images that might violate copyright laws and individual privacy. This phenomenon is termed imitation -- generation of images with content that has recognizable similarity to its training images. In this work we estimate the point at which a model was trained on enough instances of a concept to be able to imitate it -- the imitation threshold. We posit this question as a new problem and propose an efficient approach that estimates the imitation threshold without incurring the colossal cost of training these models from scratch. We experiment with two domains -- human faces and art styles, and evaluate four text-to-image models that were trained on three pretraining datasets. We estimate the imitation threshold of these models to be in the range of 200-700 images, depending on the domain and the model. The imitation threshold provides an empirical basis for copyright violation claims and acts as a guiding principle for text-to-image model developers that aim to comply with copyright and privacy laws. Website: https://how-many-van-goghs-does-it-take.github.io/. Code: https://github.com/vsahil/MIMETIC-2.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15002
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Many Images Does It Take? Estimating Imitation Thresholds in Text-to-Image Models
Verma, Sahil
Rassin, Royi
Das, Arnav
Bhatt, Gantavya
Seshadri, Preethi
Shah, Chirag
Bilmes, Jeff
Hajishirzi, Hannaneh
Elazar, Yanai
Computer Vision and Pattern Recognition
Text-to-image models are trained using large datasets of image-text pairs collected from the internet. These datasets often include copyrighted and private images. Training models on such datasets enables them to generate images that might violate copyright laws and individual privacy. This phenomenon is termed imitation -- generation of images with content that has recognizable similarity to its training images. In this work we estimate the point at which a model was trained on enough instances of a concept to be able to imitate it -- the imitation threshold. We posit this question as a new problem and propose an efficient approach that estimates the imitation threshold without incurring the colossal cost of training these models from scratch. We experiment with two domains -- human faces and art styles, and evaluate four text-to-image models that were trained on three pretraining datasets. We estimate the imitation threshold of these models to be in the range of 200-700 images, depending on the domain and the model. The imitation threshold provides an empirical basis for copyright violation claims and acts as a guiding principle for text-to-image model developers that aim to comply with copyright and privacy laws. Website: https://how-many-van-goghs-does-it-take.github.io/. Code: https://github.com/vsahil/MIMETIC-2.
title How Many Images Does It Take? Estimating Imitation Thresholds in Text-to-Image Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.15002