FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918319322300416 |
|---|---|
| author | FSVideo Team Chen, Qingyu Fang, Zhiyuan Huang, Haibin Huang, Xinwei Jin, Tong Lin, Minxuan Liu, Bo Liu, Celong Ma, Chongyang Mei, Xing Shen, Xiaohui Shen, Yaojie Tan, Fuwen Wang, Angtian Yang, Xiao Yang, Yiding Yuan, Jiamin Zhang, Lingxi Zhang, Yuxin |
| author_facet | FSVideo Team Chen, Qingyu Fang, Zhiyuan Huang, Haibin Huang, Xinwei Jin, Tong Lin, Minxuan Liu, Bo Liu, Celong Ma, Chongyang Mei, Xing Shen, Xiaohui Shen, Yaojie Tan, Fuwen Wang, Angtian Yang, Xiao Yang, Yiding Yuan, Jiamin Zhang, Lingxi Zhang, Yuxin |
| contents | We introduce FSVideo, a fast speed transformer-based image-to-video (I2V) diffusion framework. We build our framework on the following key components: 1.) a new video autoencoder with highly-compressed latent space ($64\times64\times4$ spatial-temporal downsampling ratio), achieving competitive reconstruction quality; 2.) a diffusion transformer (DIT) architecture with a new layer memory design to enhance inter-layer information flow and context reuse within DIT, and 3.) a multi-resolution generation strategy via a few-step DIT upsampler to increase video fidelity. Our final model, which contains a 14B DIT base model and a 14B DIT upsampler, achieves competitive performance against other popular open-source models, while being an order of magnitude faster. We discuss our model design as well as training strategies in this report. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_02092 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space FSVideo Team Chen, Qingyu Fang, Zhiyuan Huang, Haibin Huang, Xinwei Jin, Tong Lin, Minxuan Liu, Bo Liu, Celong Ma, Chongyang Mei, Xing Shen, Xiaohui Shen, Yaojie Tan, Fuwen Wang, Angtian Yang, Xiao Yang, Yiding Yuan, Jiamin Zhang, Lingxi Zhang, Yuxin Computer Vision and Pattern Recognition We introduce FSVideo, a fast speed transformer-based image-to-video (I2V) diffusion framework. We build our framework on the following key components: 1.) a new video autoencoder with highly-compressed latent space ($64\times64\times4$ spatial-temporal downsampling ratio), achieving competitive reconstruction quality; 2.) a diffusion transformer (DIT) architecture with a new layer memory design to enhance inter-layer information flow and context reuse within DIT, and 3.) a multi-resolution generation strategy via a few-step DIT upsampler to increase video fidelity. Our final model, which contains a 14B DIT base model and a 14B DIT upsampler, achieves competitive performance against other popular open-source models, while being an order of magnitude faster. We discuss our model design as well as training strategies in this report. |
| title | FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2602.02092 |