BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Jiacheng, Mehran, Ramin, Jia, Xuhui, Xie, Saining, Woo, Sanghyun
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912565319172096
author Chen, Jiacheng
Mehran, Ramin
Jia, Xuhui
Xie, Saining
Woo, Sanghyun
author_facet Chen, Jiacheng
Mehran, Ramin
Jia, Xuhui
Xie, Saining
Woo, Sanghyun
contents We present BlenderFusion, a generative visual compositing framework that synthesizes new scenes by recomposing objects, camera, and background. It follows a layering-editing-compositing pipeline: (i) segmenting and converting visual inputs into editable 3D entities (layering), (ii) editing them in Blender with 3D-grounded control (editing), and (iii) fusing them into a coherent scene using a generative compositor (compositing). Our generative compositor extends a pre-trained diffusion model to process both the original (source) and edited (target) scenes in parallel. It is fine-tuned on video frames with two key training strategies: (i) source masking, enabling flexible modifications like background replacement; (ii) simulated object jittering, facilitating disentangled control over objects and camera. BlenderFusion significantly outperforms prior methods in complex compositional scene editing tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17450
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
Chen, Jiacheng
Mehran, Ramin
Jia, Xuhui
Xie, Saining
Woo, Sanghyun
Computer Vision and Pattern Recognition
Graphics
We present BlenderFusion, a generative visual compositing framework that synthesizes new scenes by recomposing objects, camera, and background. It follows a layering-editing-compositing pipeline: (i) segmenting and converting visual inputs into editable 3D entities (layering), (ii) editing them in Blender with 3D-grounded control (editing), and (iii) fusing them into a coherent scene using a generative compositor (compositing). Our generative compositor extends a pre-trained diffusion model to process both the original (source) and edited (target) scenes in parallel. It is fine-tuned on video frames with two key training strategies: (i) source masking, enabling flexible modifications like background replacement; (ii) simulated object jittering, facilitating disentangled control over objects and camera. BlenderFusion significantly outperforms prior methods in complex compositional scene editing tasks.
title BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2506.17450