SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xinyu, Qian, Yuyi, Lin, Jiang, Wang, Shenyi, Wang, Gao, Zhang, Zhiqiu, Zhang, Jizhi, Wang, Mingjie, Tang, Qiang, Wang, Qian, Wu, Song, Yi, Zili
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913154912485376
author Chen, Xinyu
Qian, Yuyi
Lin, Jiang
Wang, Shenyi
Wang, Gao
Zhang, Zhiqiu
Zhang, Jizhi
Wang, Mingjie
Tang, Qiang
Wang, Qian
Wu, Song
Yi, Zili
author_facet Chen, Xinyu
Qian, Yuyi
Lin, Jiang
Wang, Shenyi
Wang, Gao
Zhang, Zhiqiu
Zhang, Jizhi
Wang, Mingjie
Tang, Qiang
Wang, Qian
Wu, Song
Yi, Zili
contents Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, current approaches are often hindered by a reliance on explicit motion engineering or resource-intensive retraining, restricting their flexibility and generalization. To bridge this gap, we present \textit{SimInsert}, a training-free paradigm that efficiently decouples the task into intuitive single-frame editing and semantic motion description. By harnessing the robust generative priors of image-to-video diffusion models, SimInsert propagates edits temporally, strictly preserving background invariance while enabling plausible, text-driven interactions between the inserted object and the dynamic environment. Our approach hinges on non-invasive guidance mechanisms that enforce structural consistency, facilitate seamless boundary fusion, and counteract the fidelity drift that typically accumulates during the denoising trajectory. Extensive quantitative experiments validate our efficacy: SimInsert surpasses state-of-the-art methods with an 18.8\% gain in PSNR, 20.1\% in SSIM, and a 44.1\% decrease in LPIPS, offering a streamlined solution for high-fidelity video editing.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23245
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion
Chen, Xinyu
Qian, Yuyi
Lin, Jiang
Wang, Shenyi
Wang, Gao
Zhang, Zhiqiu
Zhang, Jizhi
Wang, Mingjie
Tang, Qiang
Wang, Qian
Wu, Song
Yi, Zili
Computer Vision and Pattern Recognition
Artificial Intelligence
Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, current approaches are often hindered by a reliance on explicit motion engineering or resource-intensive retraining, restricting their flexibility and generalization. To bridge this gap, we present \textit{SimInsert}, a training-free paradigm that efficiently decouples the task into intuitive single-frame editing and semantic motion description. By harnessing the robust generative priors of image-to-video diffusion models, SimInsert propagates edits temporally, strictly preserving background invariance while enabling plausible, text-driven interactions between the inserted object and the dynamic environment. Our approach hinges on non-invasive guidance mechanisms that enforce structural consistency, facilitate seamless boundary fusion, and counteract the fidelity drift that typically accumulates during the denoising trajectory. Extensive quantitative experiments validate our efficacy: SimInsert surpasses state-of-the-art methods with an 18.8\% gain in PSNR, 20.1\% in SSIM, and a 44.1\% decrease in LPIPS, offering a streamlined solution for high-fidelity video editing.
title SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.23245