BOOTPLACE: Bootstrapped Object Placement with Detection Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Hang, Zuo, Xinxin, Ma, Rui, Cheng, Li
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908288251068416
author Zhou, Hang
Zuo, Xinxin
Ma, Rui
Cheng, Li
author_facet Zhou, Hang
Zuo, Xinxin
Ma, Rui
Cheng, Li
contents In this paper, we tackle the copy-paste image-to-image composition problem with a focus on object placement learning. Prior methods have leveraged generative models to reduce the reliance for dense supervision. However, this often limits their capacity to model complex data distributions. Alternatively, transformer networks with a sparse contrastive loss have been explored, but their over-relaxed regularization often leads to imprecise object placement. We introduce BOOTPLACE, a novel paradigm that formulates object placement as a placement-by-detection problem. Our approach begins by identifying suitable regions of interest for object placement. This is achieved by training a specialized detection transformer on object-subtracted backgrounds, enhanced with multi-object supervisions. It then semantically associates each target compositing object with detected regions based on their complementary characteristics. Through a boostrapped training approach applied to randomly object-subtracted images, our model enforces meaningful placements through extensive paired data augmentation. Experimental results on established benchmarks demonstrate BOOTPLACE's superior performance in object repositioning, markedly surpassing state-of-the-art baselines on Cityscapes and OPA datasets with notable improvements in IOU scores. Additional ablation studies further showcase the compositionality and generalizability of our approach, supported by user study evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21991
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BOOTPLACE: Bootstrapped Object Placement with Detection Transformers
Zhou, Hang
Zuo, Xinxin
Ma, Rui
Cheng, Li
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
In this paper, we tackle the copy-paste image-to-image composition problem with a focus on object placement learning. Prior methods have leveraged generative models to reduce the reliance for dense supervision. However, this often limits their capacity to model complex data distributions. Alternatively, transformer networks with a sparse contrastive loss have been explored, but their over-relaxed regularization often leads to imprecise object placement. We introduce BOOTPLACE, a novel paradigm that formulates object placement as a placement-by-detection problem. Our approach begins by identifying suitable regions of interest for object placement. This is achieved by training a specialized detection transformer on object-subtracted backgrounds, enhanced with multi-object supervisions. It then semantically associates each target compositing object with detected regions based on their complementary characteristics. Through a boostrapped training approach applied to randomly object-subtracted images, our model enforces meaningful placements through extensive paired data augmentation. Experimental results on established benchmarks demonstrate BOOTPLACE's superior performance in object repositioning, markedly surpassing state-of-the-art baselines on Cityscapes and OPA datasets with notable improvements in IOU scores. Additional ablation studies further showcase the compositionality and generalizability of our approach, supported by user study evaluations.
title BOOTPLACE: Bootstrapped Object Placement with Detection Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
url https://arxiv.org/abs/2503.21991