MObI: Multimodal Object Inpainting Using Diffusion Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Buburuzan, Alexandru, Sharma, Anuj, Redford, John, Dokania, Puneet K., Mueller, Romain
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916700786524160
author Buburuzan, Alexandru
Sharma, Anuj
Redford, John
Dokania, Puneet K.
Mueller, Romain
author_facet Buburuzan, Alexandru
Sharma, Anuj
Redford, John
Dokania, Puneet K.
Mueller, Romain
contents Safety-critical applications, such as autonomous driving, require extensive multimodal data for rigorous testing. Methods based on synthetic data are gaining prominence due to the cost and complexity of gathering real-world data but require a high degree of realism and controllability in order to be useful. This paper introduces MObI, a novel framework for Multimodal Object Inpainting that leverages a diffusion model to create realistic and controllable object inpaintings across perceptual modalities, demonstrated for both camera and lidar simultaneously. Using a single reference RGB image, MObI enables objects to be seamlessly inserted into existing multimodal scenes at a 3D location specified by a bounding box, while maintaining semantic consistency and multimodal coherence. Unlike traditional inpainting methods that rely solely on edit masks, our 3D bounding box conditioning gives objects accurate spatial positioning and realistic scaling. As a result, our approach can be used to insert novel objects flexibly into multimodal scenes, providing significant advantages for testing perception models.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03173
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MObI: Multimodal Object Inpainting Using Diffusion Models
Buburuzan, Alexandru
Sharma, Anuj
Redford, John
Dokania, Puneet K.
Mueller, Romain
Computer Vision and Pattern Recognition
Safety-critical applications, such as autonomous driving, require extensive multimodal data for rigorous testing. Methods based on synthetic data are gaining prominence due to the cost and complexity of gathering real-world data but require a high degree of realism and controllability in order to be useful. This paper introduces MObI, a novel framework for Multimodal Object Inpainting that leverages a diffusion model to create realistic and controllable object inpaintings across perceptual modalities, demonstrated for both camera and lidar simultaneously. Using a single reference RGB image, MObI enables objects to be seamlessly inserted into existing multimodal scenes at a 3D location specified by a bounding box, while maintaining semantic consistency and multimodal coherence. Unlike traditional inpainting methods that rely solely on edit masks, our 3D bounding box conditioning gives objects accurate spatial positioning and realistic scaling. As a result, our approach can be used to insert novel objects flexibly into multimodal scenes, providing significant advantages for testing perception models.
title MObI: Multimodal Object Inpainting Using Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.03173