DeCLIP: Decoding CLIP representations for deepfake localization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Smeu, Stefan, Oneata, Elisabeta, Oneata, Dan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912150429106176
author Smeu, Stefan
Oneata, Elisabeta
Oneata, Dan
author_facet Smeu, Stefan
Oneata, Elisabeta
Oneata, Dan
contents Generative models can create entirely new images, but they can also partially modify real images in ways that are undetectable to the human eye. In this paper, we address the challenge of automatically detecting such local manipulations. One of the most pressing problems in deepfake detection remains the ability of models to generalize to different classes of generators. In the case of fully manipulated images, representations extracted from large self-supervised models (such as CLIP) provide a promising direction towards more robust detectors. Here, we introduce DeCLIP, a first attempt to leverage such large pretrained features for detecting local manipulations. We show that, when combined with a reasonably large convolutional decoder, pretrained self-supervised representations are able to perform localization and improve generalization capabilities over existing methods. Unlike previous work, our approach is able to perform localization on the challenging case of latent diffusion models, where the entire image is affected by the fingerprint of the generator. Moreover, we observe that this type of data, which combines local semantic information with a global fingerprint, provides more stable generalization than other categories of generative methods.
format Preprint
id arxiv_https___arxiv_org_abs_2409_08849
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DeCLIP: Decoding CLIP representations for deepfake localization
Smeu, Stefan
Oneata, Elisabeta
Oneata, Dan
Computer Vision and Pattern Recognition
Machine Learning
Generative models can create entirely new images, but they can also partially modify real images in ways that are undetectable to the human eye. In this paper, we address the challenge of automatically detecting such local manipulations. One of the most pressing problems in deepfake detection remains the ability of models to generalize to different classes of generators. In the case of fully manipulated images, representations extracted from large self-supervised models (such as CLIP) provide a promising direction towards more robust detectors. Here, we introduce DeCLIP, a first attempt to leverage such large pretrained features for detecting local manipulations. We show that, when combined with a reasonably large convolutional decoder, pretrained self-supervised representations are able to perform localization and improve generalization capabilities over existing methods. Unlike previous work, our approach is able to perform localization on the challenging case of latent diffusion models, where the entire image is affected by the fingerprint of the generator. Moreover, we observe that this type of data, which combines local semantic information with a global fingerprint, provides more stable generalization than other categories of generative methods.
title DeCLIP: Decoding CLIP representations for deepfake localization
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2409.08849