InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Hoiyeong, Jang, Hyojin, Kim, Jeongho, Hyung, Junha, Kim, Kinam, Kim, Dongjin, Choi, Huijin, Kim, Hyeonji, Choo, Jaegul
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908722387746816
author Jin, Hoiyeong
Jang, Hyojin
Kim, Jeongho
Hyung, Junha
Kim, Kinam
Kim, Dongjin
Choi, Huijin
Kim, Hyeonji
Choo, Jaegul
author_facet Jin, Hoiyeong
Jang, Hyojin
Kim, Jeongho
Hyung, Junha
Kim, Kinam
Kim, Dongjin
Choi, Huijin
Kim, Hyeonji
Choo, Jaegul
contents Recent advances in diffusion-based video generation have opened new possibilities for controllable video editing, yet realistic video object insertion (VOI) remains challenging due to limited 4D scene understanding and inadequate handling of occlusion and lighting effects. We present InsertAnywhere, a new VOI framework that achieves geometrically consistent object placement and appearance-faithful video synthesis. Our method begins with a 4D aware mask generation module that reconstructs the scene geometry and propagates user specified object placement across frames while maintaining temporal coherence and occlusion consistency. Building upon this spatial foundation, we extend a diffusion based video generation model to jointly synthesize the inserted object and its surrounding local variations such as illumination and shading. To enable supervised training, we introduce ROSE++, an illumination aware synthetic dataset constructed by transforming the ROSE object removal dataset into triplets of object removed video, object present video, and a VLM generated reference image. Through extensive experiments, we demonstrate that our framework produces geometrically plausible and visually coherent object insertions across diverse real world scenarios, significantly outperforming existing research and commercial models.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17504
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
Jin, Hoiyeong
Jang, Hyojin
Kim, Jeongho
Hyung, Junha
Kim, Kinam
Kim, Dongjin
Choi, Huijin
Kim, Hyeonji
Choo, Jaegul
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advances in diffusion-based video generation have opened new possibilities for controllable video editing, yet realistic video object insertion (VOI) remains challenging due to limited 4D scene understanding and inadequate handling of occlusion and lighting effects. We present InsertAnywhere, a new VOI framework that achieves geometrically consistent object placement and appearance-faithful video synthesis. Our method begins with a 4D aware mask generation module that reconstructs the scene geometry and propagates user specified object placement across frames while maintaining temporal coherence and occlusion consistency. Building upon this spatial foundation, we extend a diffusion based video generation model to jointly synthesize the inserted object and its surrounding local variations such as illumination and shading. To enable supervised training, we introduce ROSE++, an illumination aware synthetic dataset constructed by transforming the ROSE object removal dataset into triplets of object removed video, object present video, and a VLM generated reference image. Through extensive experiments, we demonstrate that our framework produces geometrically plausible and visually coherent object insertions across diverse real world scenarios, significantly outperforming existing research and commercial models.
title InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.17504