ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Ruonan, Tan, Zhenxiong, Chen, Zigeng, Liu, Songhua, Wang, Xinchao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915867278704640
author Yu, Ruonan
Tan, Zhenxiong
Chen, Zigeng
Liu, Songhua
Wang, Xinchao
author_facet Yu, Ruonan
Tan, Zhenxiong
Chen, Zigeng
Liu, Songhua
Wang, Xinchao
contents Diffusion Transformers (DiTs) have demonstrated remarkable scalability and quality in image and video generation, prompting growing interest in extending them to controllable generation and editing tasks. However, compared to the image counterparts, progress in video control and editing remains limited, mainly due to the scarcity of paired video data and the high computational cost of training video diffusion models. To address this issue, in this paper, we propose a video-free tuning framework termed ViFeEdit for video diffusion transformers. Without requiring any forms of video training data, ViFeEdit achieves versatile video generation and editing, adapted solely with 2D images. At the core of our approach is an architectural reparameterization that decouples spatial independence from the full 3D attention in modern video diffusion transformers, which enables visually faithful editing while maintaining temporal consistency with only minimal additional parameters. Moreover, this design operates in a dual-path pipeline with separate timestep embeddings for noise scheduling, exhibiting strong adaptability to diverse conditioning signals. Extensive experiments demonstrate that our method delivers promising results of controllable video generation and editing with only minimal training on 2D image data. Codes are available https://github.com/Lexie-YU/ViFeEdit.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15478
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer
Yu, Ruonan
Tan, Zhenxiong
Chen, Zigeng
Liu, Songhua
Wang, Xinchao
Computer Vision and Pattern Recognition
Diffusion Transformers (DiTs) have demonstrated remarkable scalability and quality in image and video generation, prompting growing interest in extending them to controllable generation and editing tasks. However, compared to the image counterparts, progress in video control and editing remains limited, mainly due to the scarcity of paired video data and the high computational cost of training video diffusion models. To address this issue, in this paper, we propose a video-free tuning framework termed ViFeEdit for video diffusion transformers. Without requiring any forms of video training data, ViFeEdit achieves versatile video generation and editing, adapted solely with 2D images. At the core of our approach is an architectural reparameterization that decouples spatial independence from the full 3D attention in modern video diffusion transformers, which enables visually faithful editing while maintaining temporal consistency with only minimal additional parameters. Moreover, this design operates in a dual-path pipeline with separate timestep embeddings for noise scheduling, exhibiting strong adaptability to diverse conditioning signals. Extensive experiments demonstrate that our method delivers promising results of controllable video generation and editing with only minimal training on 2D image data. Codes are available https://github.com/Lexie-YU/ViFeEdit.
title ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.15478