Magic Insert: Style-Aware Drag-and-Drop
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866910510198292480 |
|---|---|
| author | Ruiz, Nataniel Li, Yuanzhen Wadhwa, Neal Pritch, Yael Rubinstein, Michael Jacobs, David E. Fruchter, Shlomi |
| author_facet | Ruiz, Nataniel Li, Yuanzhen Wadhwa, Neal Pritch, Yael Rubinstein, Michael Jacobs, David E. Fruchter, Shlomi |
| contents | We present Magic Insert, a method for dragging-and-dropping subjects from a user-provided image into a target image of a different style in a physically plausible manner while matching the style of the target image. This work formalizes the problem of style-aware drag-and-drop and presents a method for tackling it by addressing two sub-problems: style-aware personalization and realistic object insertion in stylized images. For style-aware personalization, our method first fine-tunes a pretrained text-to-image diffusion model using LoRA and learned text tokens on the subject image, and then infuses it with a CLIP representation of the target style. For object insertion, we use Bootstrapped Domain Adaption to adapt a domain-specific photorealistic object insertion model to the domain of diverse artistic styles. Overall, the method significantly outperforms traditional approaches such as inpainting. Finally, we present a dataset, SubjectPlop, to facilitate evaluation and future progress in this area. Project page: https://magicinsert.github.io/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_02489 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Magic Insert: Style-Aware Drag-and-Drop Ruiz, Nataniel Li, Yuanzhen Wadhwa, Neal Pritch, Yael Rubinstein, Michael Jacobs, David E. Fruchter, Shlomi Computer Vision and Pattern Recognition Artificial Intelligence Graphics Human-Computer Interaction Machine Learning We present Magic Insert, a method for dragging-and-dropping subjects from a user-provided image into a target image of a different style in a physically plausible manner while matching the style of the target image. This work formalizes the problem of style-aware drag-and-drop and presents a method for tackling it by addressing two sub-problems: style-aware personalization and realistic object insertion in stylized images. For style-aware personalization, our method first fine-tunes a pretrained text-to-image diffusion model using LoRA and learned text tokens on the subject image, and then infuses it with a CLIP representation of the target style. For object insertion, we use Bootstrapped Domain Adaption to adapt a domain-specific photorealistic object insertion model to the domain of diverse artistic styles. Overall, the method significantly outperforms traditional approaches such as inpainting. Finally, we present a dataset, SubjectPlop, to facilitate evaluation and future progress in this area. Project page: https://magicinsert.github.io/ |
| title | Magic Insert: Style-Aware Drag-and-Drop |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Graphics Human-Computer Interaction Machine Learning |
| url | https://arxiv.org/abs/2407.02489 |