Repositioning the Subject within Image

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Yikai, Cao, Chenjie, Fan, Ke, Dong, Qiaole, Li, Yifan, Xue, Xiangyang, Fu, Yanwei
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912125495017472
author Wang, Yikai
Cao, Chenjie
Fan, Ke
Dong, Qiaole
Li, Yifan
Xue, Xiangyang
Fu, Yanwei
author_facet Wang, Yikai
Cao, Chenjie
Fan, Ke
Dong, Qiaole
Li, Yifan
Xue, Xiangyang
Fu, Yanwei
contents Current image manipulation primarily centers on static manipulation, such as replacing specific regions within an image or altering its overall style. In this paper, we introduce an innovative dynamic manipulation task, subject repositioning. This task involves relocating a user-specified subject to a desired position while preserving the image's fidelity. Our research reveals that the fundamental sub-tasks of subject repositioning, which include filling the void left by the repositioned subject, reconstructing obscured portions of the subject and blending the subject to be consistent with surrounding areas, can be effectively reformulated as a unified, prompt-guided inpainting task. Consequently, we can employ a single diffusion generative model to address these sub-tasks using various task prompts learned through our proposed task inversion technique. Additionally, we integrate pre-processing and post-processing techniques to further enhance the quality of subject repositioning. These elements together form our SEgment-gEnerate-and-bLEnd (SEELE) framework. To assess SEELE's effectiveness in subject repositioning, we assemble a real-world subject repositioning dataset called ReS. Results of SEELE on ReS demonstrate its efficacy. Code and ReS dataset are available at https://yikai-wang.github.io/seele/.
format Preprint
id arxiv_https___arxiv_org_abs_2401_16861
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Repositioning the Subject within Image
Wang, Yikai
Cao, Chenjie
Fan, Ke
Dong, Qiaole
Li, Yifan
Xue, Xiangyang
Fu, Yanwei
Computer Vision and Pattern Recognition
Current image manipulation primarily centers on static manipulation, such as replacing specific regions within an image or altering its overall style. In this paper, we introduce an innovative dynamic manipulation task, subject repositioning. This task involves relocating a user-specified subject to a desired position while preserving the image's fidelity. Our research reveals that the fundamental sub-tasks of subject repositioning, which include filling the void left by the repositioned subject, reconstructing obscured portions of the subject and blending the subject to be consistent with surrounding areas, can be effectively reformulated as a unified, prompt-guided inpainting task. Consequently, we can employ a single diffusion generative model to address these sub-tasks using various task prompts learned through our proposed task inversion technique. Additionally, we integrate pre-processing and post-processing techniques to further enhance the quality of subject repositioning. These elements together form our SEgment-gEnerate-and-bLEnd (SEELE) framework. To assess SEELE's effectiveness in subject repositioning, we assemble a real-world subject repositioning dataset called ReS. Results of SEELE on ReS demonstrate its efficacy. Code and ReS dataset are available at https://yikai-wang.github.io/seele/.
title Repositioning the Subject within Image
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.16861