OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jinshu, Li, Xinghui, Bai, Xu, Ma, Tianxiang, Zhang, Pengze, Chen, Zhuowei, Li, Gen, Liu, Lijie, Zhao, Songtao, Li, Bingchuan, He, Qian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918145580597248
author Chen, Jinshu
Li, Xinghui
Bai, Xu
Ma, Tianxiang
Zhang, Pengze
Chen, Zhuowei
Li, Gen
Liu, Lijie
Zhao, Songtao
Li, Bingchuan
He, Qian
author_facet Chen, Jinshu
Li, Xinghui
Bai, Xu
Ma, Tianxiang
Zhang, Pengze
Chen, Zhuowei
Li, Gen
Liu, Lijie
Zhao, Songtao
Li, Bingchuan
He, Qian
contents Recent advances in video insertion based on diffusion models are impressive. However, existing methods rely on complex control signals but struggle with subject consistency, limiting their practical applicability. In this paper, we focus on the task of Mask-free Video Insertion and aim to resolve three key challenges: data scarcity, subject-scene equilibrium, and insertion harmonization. To address the data scarcity, we propose a new data pipeline InsertPipe, constructing diverse cross-pair data automatically. Building upon our data pipeline, we develop OmniInsert, a novel unified framework for mask-free video insertion from both single and multiple subject references. Specifically, to maintain subject-scene equilibrium, we introduce a simple yet effective Condition-Specific Feature Injection mechanism to distinctly inject multi-source conditions and propose a novel Progressive Training strategy that enables the model to balance feature injection from subjects and source video. Meanwhile, we design the Subject-Focused Loss to improve the detailed appearance of the subjects. To further enhance insertion harmonization, we propose an Insertive Preference Optimization methodology to optimize the model by simulating human preferences, and incorporate a Context-Aware Rephraser module during reference to seamlessly integrate the subject into the original scenes. To address the lack of a benchmark for the field, we introduce InsertBench, a comprehensive benchmark comprising diverse scenes with meticulously selected subjects. Evaluation on InsertBench indicates OmniInsert outperforms state-of-the-art closed-source commercial solutions. The code will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17627
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
Chen, Jinshu
Li, Xinghui
Bai, Xu
Ma, Tianxiang
Zhang, Pengze
Chen, Zhuowei
Li, Gen
Liu, Lijie
Zhao, Songtao
Li, Bingchuan
He, Qian
Computer Vision and Pattern Recognition
Recent advances in video insertion based on diffusion models are impressive. However, existing methods rely on complex control signals but struggle with subject consistency, limiting their practical applicability. In this paper, we focus on the task of Mask-free Video Insertion and aim to resolve three key challenges: data scarcity, subject-scene equilibrium, and insertion harmonization. To address the data scarcity, we propose a new data pipeline InsertPipe, constructing diverse cross-pair data automatically. Building upon our data pipeline, we develop OmniInsert, a novel unified framework for mask-free video insertion from both single and multiple subject references. Specifically, to maintain subject-scene equilibrium, we introduce a simple yet effective Condition-Specific Feature Injection mechanism to distinctly inject multi-source conditions and propose a novel Progressive Training strategy that enables the model to balance feature injection from subjects and source video. Meanwhile, we design the Subject-Focused Loss to improve the detailed appearance of the subjects. To further enhance insertion harmonization, we propose an Insertive Preference Optimization methodology to optimize the model by simulating human preferences, and incorporate a Context-Aware Rephraser module during reference to seamlessly integrate the subject into the original scenes. To address the lack of a benchmark for the field, we introduce InsertBench, a comprehensive benchmark comprising diverse scenes with meticulously selected subjects. Evaluation on InsertBench indicates OmniInsert outperforms state-of-the-art closed-source commercial solutions. The code will be released.
title OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.17627