VFXMaster: Unlocking Dynamic Visual Effect Generation via In-Context Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Baolu, Zhang, Yiming, Wang, Qinghe, Ma, Liqian, Shi, Xiaoyu, Wang, Xintao, Wan, Pengfei, Yin, Zhenfei, Zhuge, Yunzhi, Lu, Huchuan, Jia, Xu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918177902952448
author Li, Baolu
Zhang, Yiming
Wang, Qinghe
Ma, Liqian
Shi, Xiaoyu
Wang, Xintao
Wan, Pengfei
Yin, Zhenfei
Zhuge, Yunzhi
Lu, Huchuan
Jia, Xu
author_facet Li, Baolu
Zhang, Yiming
Wang, Qinghe
Ma, Liqian
Shi, Xiaoyu
Wang, Xintao
Wan, Pengfei
Yin, Zhenfei
Zhuge, Yunzhi
Lu, Huchuan
Jia, Xu
contents Visual effects (VFX) are crucial to the expressive power of digital media, yet their creation remains a major challenge for generative AI. Prevailing methods often rely on the one-LoRA-per-effect paradigm, which is resource-intensive and fundamentally incapable of generalizing to unseen effects, thus limiting scalability and creation. To address this challenge, we introduce VFXMaster, the first unified, reference-based framework for VFX video generation. It recasts effect generation as an in-context learning task, enabling it to reproduce diverse dynamic effects from a reference video onto target content. In addition, it demonstrates remarkable generalization to unseen effect categories. Specifically, we design an in-context conditioning strategy that prompts the model with a reference example. An in-context attention mask is designed to precisely decouple and inject the essential effect attributes, allowing a single unified model to master the effect imitation without information leakage. In addition, we propose an efficient one-shot effect adaptation mechanism to boost generalization capability on tough unseen effects from a single user-provided video rapidly. Extensive experiments demonstrate that our method effectively imitates various categories of effect information and exhibits outstanding generalization to out-of-domain effects. To foster future research, we will release our code, models, and a comprehensive dataset to the community.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25772
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VFXMaster: Unlocking Dynamic Visual Effect Generation via In-Context Learning
Li, Baolu
Zhang, Yiming
Wang, Qinghe
Ma, Liqian
Shi, Xiaoyu
Wang, Xintao
Wan, Pengfei
Yin, Zhenfei
Zhuge, Yunzhi
Lu, Huchuan
Jia, Xu
Computer Vision and Pattern Recognition
Visual effects (VFX) are crucial to the expressive power of digital media, yet their creation remains a major challenge for generative AI. Prevailing methods often rely on the one-LoRA-per-effect paradigm, which is resource-intensive and fundamentally incapable of generalizing to unseen effects, thus limiting scalability and creation. To address this challenge, we introduce VFXMaster, the first unified, reference-based framework for VFX video generation. It recasts effect generation as an in-context learning task, enabling it to reproduce diverse dynamic effects from a reference video onto target content. In addition, it demonstrates remarkable generalization to unseen effect categories. Specifically, we design an in-context conditioning strategy that prompts the model with a reference example. An in-context attention mask is designed to precisely decouple and inject the essential effect attributes, allowing a single unified model to master the effect imitation without information leakage. In addition, we propose an efficient one-shot effect adaptation mechanism to boost generalization capability on tough unseen effects from a single user-provided video rapidly. Extensive experiments demonstrate that our method effectively imitates various categories of effect information and exhibits outstanding generalization to out-of-domain effects. To foster future research, we will release our code, models, and a comprehensive dataset to the community.
title VFXMaster: Unlocking Dynamic Visual Effect Generation via In-Context Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.25772