FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: FSVideo Team, Chen, Qingyu, Fang, Zhiyuan, Huang, Haibin, Huang, Xinwei, Jin, Tong, Lin, Minxuan, Liu, Bo, Liu, Celong, Ma, Chongyang, Mei, Xing, Shen, Xiaohui, Shen, Yaojie, Tan, Fuwen, Wang, Angtian, Yang, Xiao, Yang, Yiding, Yuan, Jiamin, Zhang, Lingxi, Zhang, Yuxin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918319322300416
author FSVideo Team
Chen, Qingyu
Fang, Zhiyuan
Huang, Haibin
Huang, Xinwei
Jin, Tong
Lin, Minxuan
Liu, Bo
Liu, Celong
Ma, Chongyang
Mei, Xing
Shen, Xiaohui
Shen, Yaojie
Tan, Fuwen
Wang, Angtian
Yang, Xiao
Yang, Yiding
Yuan, Jiamin
Zhang, Lingxi
Zhang, Yuxin
author_facet FSVideo Team
Chen, Qingyu
Fang, Zhiyuan
Huang, Haibin
Huang, Xinwei
Jin, Tong
Lin, Minxuan
Liu, Bo
Liu, Celong
Ma, Chongyang
Mei, Xing
Shen, Xiaohui
Shen, Yaojie
Tan, Fuwen
Wang, Angtian
Yang, Xiao
Yang, Yiding
Yuan, Jiamin
Zhang, Lingxi
Zhang, Yuxin
contents We introduce FSVideo, a fast speed transformer-based image-to-video (I2V) diffusion framework. We build our framework on the following key components: 1.) a new video autoencoder with highly-compressed latent space ($64\times64\times4$ spatial-temporal downsampling ratio), achieving competitive reconstruction quality; 2.) a diffusion transformer (DIT) architecture with a new layer memory design to enhance inter-layer information flow and context reuse within DIT, and 3.) a multi-resolution generation strategy via a few-step DIT upsampler to increase video fidelity. Our final model, which contains a 14B DIT base model and a 14B DIT upsampler, achieves competitive performance against other popular open-source models, while being an order of magnitude faster. We discuss our model design as well as training strategies in this report.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02092
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space
FSVideo Team
Chen, Qingyu
Fang, Zhiyuan
Huang, Haibin
Huang, Xinwei
Jin, Tong
Lin, Minxuan
Liu, Bo
Liu, Celong
Ma, Chongyang
Mei, Xing
Shen, Xiaohui
Shen, Yaojie
Tan, Fuwen
Wang, Angtian
Yang, Xiao
Yang, Yiding
Yuan, Jiamin
Zhang, Lingxi
Zhang, Yuxin
Computer Vision and Pattern Recognition
We introduce FSVideo, a fast speed transformer-based image-to-video (I2V) diffusion framework. We build our framework on the following key components: 1.) a new video autoencoder with highly-compressed latent space ($64\times64\times4$ spatial-temporal downsampling ratio), achieving competitive reconstruction quality; 2.) a diffusion transformer (DIT) architecture with a new layer memory design to enhance inter-layer information flow and context reuse within DIT, and 3.) a multi-resolution generation strategy via a few-step DIT upsampler to increase video fidelity. Our final model, which contains a 14B DIT base model and a 14B DIT upsampler, achieves competitive performance against other popular open-source models, while being an order of magnitude faster. We discuss our model design as well as training strategies in this report.
title FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.02092