Concept Lancet: Image Editing with Compositional Representation Transplant

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Luo, Jinqi, Ding, Tianjiao, Chan, Kwan Ho Ryan, Min, Hancheng, Callison-Burch, Chris, Vidal, René
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908299662721024
author Luo, Jinqi
Ding, Tianjiao
Chan, Kwan Ho Ryan
Min, Hancheng
Callison-Burch, Chris
Vidal, René
author_facet Luo, Jinqi
Ding, Tianjiao
Chan, Kwan Ho Ryan
Min, Hancheng
Callison-Burch, Chris
Vidal, René
contents Diffusion models are widely used for image editing tasks. Existing editing methods often design a representation manipulation procedure by curating an edit direction in the text embedding or score space. However, such a procedure faces a key challenge: overestimating the edit strength harms visual consistency while underestimating it fails the editing task. Notably, each source image may require a different editing strength, and it is costly to search for an appropriate strength via trial-and-error. To address this challenge, we propose Concept Lancet (CoLan), a zero-shot plug-and-play framework for principled representation manipulation in diffusion-based image editing. At inference time, we decompose the source input in the latent (text embedding or diffusion score) space as a sparse linear combination of the representations of the collected visual concepts. This allows us to accurately estimate the presence of concepts in each image, which informs the edit. Based on the editing task (replace/add/remove), we perform a customized concept transplant process to impose the corresponding editing direction. To sufficiently model the concept space, we curate a conceptual representation dataset, CoLan-150K, which contains diverse descriptions and scenarios of visual terms and phrases for the latent dictionary. Experiments on multiple diffusion-based image editing baselines show that methods equipped with CoLan achieve state-of-the-art performance in editing effectiveness and consistency preservation.
format Preprint
id arxiv_https___arxiv_org_abs_2504_02828
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Concept Lancet: Image Editing with Compositional Representation Transplant
Luo, Jinqi
Ding, Tianjiao
Chan, Kwan Ho Ryan
Min, Hancheng
Callison-Burch, Chris
Vidal, René
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Diffusion models are widely used for image editing tasks. Existing editing methods often design a representation manipulation procedure by curating an edit direction in the text embedding or score space. However, such a procedure faces a key challenge: overestimating the edit strength harms visual consistency while underestimating it fails the editing task. Notably, each source image may require a different editing strength, and it is costly to search for an appropriate strength via trial-and-error. To address this challenge, we propose Concept Lancet (CoLan), a zero-shot plug-and-play framework for principled representation manipulation in diffusion-based image editing. At inference time, we decompose the source input in the latent (text embedding or diffusion score) space as a sparse linear combination of the representations of the collected visual concepts. This allows us to accurately estimate the presence of concepts in each image, which informs the edit. Based on the editing task (replace/add/remove), we perform a customized concept transplant process to impose the corresponding editing direction. To sufficiently model the concept space, we curate a conceptual representation dataset, CoLan-150K, which contains diverse descriptions and scenarios of visual terms and phrases for the latent dictionary. Experiments on multiple diffusion-based image editing baselines show that methods equipped with CoLan achieve state-of-the-art performance in editing effectiveness and consistency preservation.
title Concept Lancet: Image Editing with Compositional Representation Transplant
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2504.02828