Magic Insert: Style-Aware Drag-and-Drop

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ruiz, Nataniel, Li, Yuanzhen, Wadhwa, Neal, Pritch, Yael, Rubinstein, Michael, Jacobs, David E., Fruchter, Shlomi
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910510198292480
author Ruiz, Nataniel
Li, Yuanzhen
Wadhwa, Neal
Pritch, Yael
Rubinstein, Michael
Jacobs, David E.
Fruchter, Shlomi
author_facet Ruiz, Nataniel
Li, Yuanzhen
Wadhwa, Neal
Pritch, Yael
Rubinstein, Michael
Jacobs, David E.
Fruchter, Shlomi
contents We present Magic Insert, a method for dragging-and-dropping subjects from a user-provided image into a target image of a different style in a physically plausible manner while matching the style of the target image. This work formalizes the problem of style-aware drag-and-drop and presents a method for tackling it by addressing two sub-problems: style-aware personalization and realistic object insertion in stylized images. For style-aware personalization, our method first fine-tunes a pretrained text-to-image diffusion model using LoRA and learned text tokens on the subject image, and then infuses it with a CLIP representation of the target style. For object insertion, we use Bootstrapped Domain Adaption to adapt a domain-specific photorealistic object insertion model to the domain of diverse artistic styles. Overall, the method significantly outperforms traditional approaches such as inpainting. Finally, we present a dataset, SubjectPlop, to facilitate evaluation and future progress in this area. Project page: https://magicinsert.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2407_02489
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Magic Insert: Style-Aware Drag-and-Drop
Ruiz, Nataniel
Li, Yuanzhen
Wadhwa, Neal
Pritch, Yael
Rubinstein, Michael
Jacobs, David E.
Fruchter, Shlomi
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Human-Computer Interaction
Machine Learning
We present Magic Insert, a method for dragging-and-dropping subjects from a user-provided image into a target image of a different style in a physically plausible manner while matching the style of the target image. This work formalizes the problem of style-aware drag-and-drop and presents a method for tackling it by addressing two sub-problems: style-aware personalization and realistic object insertion in stylized images. For style-aware personalization, our method first fine-tunes a pretrained text-to-image diffusion model using LoRA and learned text tokens on the subject image, and then infuses it with a CLIP representation of the target style. For object insertion, we use Bootstrapped Domain Adaption to adapt a domain-specific photorealistic object insertion model to the domain of diverse artistic styles. Overall, the method significantly outperforms traditional approaches such as inpainting. Finally, we present a dataset, SubjectPlop, to facilitate evaluation and future progress in this area. Project page: https://magicinsert.github.io/
title Magic Insert: Style-Aware Drag-and-Drop
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2407.02489