IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yuanhang, Song, Yiren, Bai, Junzhe, Liang, Xinran, Yang, Hu, Jin, Libiao, Mao, Qi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912771253207040
author Li, Yuanhang
Song, Yiren
Bai, Junzhe
Liang, Xinran
Yang, Hu
Jin, Libiao
Mao, Qi
author_facet Li, Yuanhang
Song, Yiren
Bai, Junzhe
Liang, Xinran
Yang, Hu
Jin, Libiao
Mao, Qi
contents We propose \textbf{IC-Effect}, an instruction-guided, DiT-based framework for few-shot video VFX editing that synthesizes complex effects (\eg flames, particles and cartoon characters) while strictly preserving spatial and temporal consistency. Video VFX editing is highly challenging because injected effects must blend seamlessly with the background, the background must remain entirely unchanged, and effect patterns must be learned efficiently from limited paired data. However, existing video editing models fail to satisfy these requirements. IC-Effect leverages the source video as clean contextual conditions, exploiting the contextual learning capability of DiT models to achieve precise background preservation and natural effect injection. A two-stage training strategy, consisting of general editing adaptation followed by effect-specific learning via Effect-LoRA, ensures strong instruction following and robust effect modeling. To further improve efficiency, we introduce spatiotemporal sparse tokenization, enabling high fidelity with substantially reduced computation. We also release a paired VFX editing dataset spanning $15$ high-quality visual styles. Extensive experiments show that IC-Effect delivers high-quality, controllable, and temporally consistent VFX editing, opening new possibilities for video creation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_15635
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning
Li, Yuanhang
Song, Yiren
Bai, Junzhe
Liang, Xinran
Yang, Hu
Jin, Libiao
Mao, Qi
Computer Vision and Pattern Recognition
Artificial Intelligence
We propose \textbf{IC-Effect}, an instruction-guided, DiT-based framework for few-shot video VFX editing that synthesizes complex effects (\eg flames, particles and cartoon characters) while strictly preserving spatial and temporal consistency. Video VFX editing is highly challenging because injected effects must blend seamlessly with the background, the background must remain entirely unchanged, and effect patterns must be learned efficiently from limited paired data. However, existing video editing models fail to satisfy these requirements. IC-Effect leverages the source video as clean contextual conditions, exploiting the contextual learning capability of DiT models to achieve precise background preservation and natural effect injection. A two-stage training strategy, consisting of general editing adaptation followed by effect-specific learning via Effect-LoRA, ensures strong instruction following and robust effect modeling. To further improve efficiency, we introduce spatiotemporal sparse tokenization, enabling high fidelity with substantially reduced computation. We also release a paired VFX editing dataset spanning $15$ high-quality visual styles. Extensive experiments show that IC-Effect delivers high-quality, controllable, and temporally consistent VFX editing, opening new possibilities for video creation.
title IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.15635