EditCLIP: Representation Learning for Image Editing

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Qian, Cvejic, Aleksandar, Eldesokey, Abdelrahman, Wonka, Peter
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912294852624384
author Wang, Qian
Cvejic, Aleksandar
Eldesokey, Abdelrahman
Wonka, Peter
author_facet Wang, Qian
Cvejic, Aleksandar
Eldesokey, Abdelrahman
Wonka, Peter
contents We introduce EditCLIP, a novel representation-learning approach for image editing. Our method learns a unified representation of edits by jointly encoding an input image and its edited counterpart, effectively capturing their transformation. To evaluate its effectiveness, we employ EditCLIP to solve two tasks: exemplar-based image editing and automated edit evaluation. In exemplar-based image editing, we replace text-based instructions in InstructPix2Pix with EditCLIP embeddings computed from a reference exemplar image pair. Experiments demonstrate that our approach outperforms state-of-the-art methods while being more efficient and versatile. For automated evaluation, EditCLIP assesses image edits by measuring the similarity between the EditCLIP embedding of a given image pair and either a textual editing instruction or the EditCLIP embedding of another reference image pair. Experiments show that EditCLIP aligns more closely with human judgments than existing CLIP-based metrics, providing a reliable measure of edit quality and structural preservation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_20318
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EditCLIP: Representation Learning for Image Editing
Wang, Qian
Cvejic, Aleksandar
Eldesokey, Abdelrahman
Wonka, Peter
Computer Vision and Pattern Recognition
We introduce EditCLIP, a novel representation-learning approach for image editing. Our method learns a unified representation of edits by jointly encoding an input image and its edited counterpart, effectively capturing their transformation. To evaluate its effectiveness, we employ EditCLIP to solve two tasks: exemplar-based image editing and automated edit evaluation. In exemplar-based image editing, we replace text-based instructions in InstructPix2Pix with EditCLIP embeddings computed from a reference exemplar image pair. Experiments demonstrate that our approach outperforms state-of-the-art methods while being more efficient and versatile. For automated evaluation, EditCLIP assesses image edits by measuring the similarity between the EditCLIP embedding of a given image pair and either a textual editing instruction or the EditCLIP embedding of another reference image pair. Experiments show that EditCLIP aligns more closely with human judgments than existing CLIP-based metrics, providing a reliable measure of edit quality and structural preservation.
title EditCLIP: Representation Learning for Image Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.20318