SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Junsong, Zhao, Yuyang, Yu, Jincheng, Chu, Ruihang, Chen, Junyu, Yang, Shuai, Wang, Xianbang, Pan, Yicheng, Zhou, Daquan, Ling, Huan, Liu, Haozhe, Yi, Hongwei, Zhang, Hao, Li, Muyang, Chen, Yukang, Cai, Han, Fidler, Sanja, Luo, Ping, Han, Song, Xie, Enze
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908588744638464
author Chen, Junsong
Zhao, Yuyang
Yu, Jincheng
Chu, Ruihang
Chen, Junyu
Yang, Shuai
Wang, Xianbang
Pan, Yicheng
Zhou, Daquan
Ling, Huan
Liu, Haozhe
Yi, Hongwei
Zhang, Hao
Li, Muyang
Chen, Yukang
Cai, Han
Fidler, Sanja
Luo, Ping
Han, Song
Xie, Enze
author_facet Chen, Junsong
Zhao, Yuyang
Yu, Jincheng
Chu, Ruihang
Chen, Junyu
Yang, Shuai
Wang, Xianbang
Pan, Yicheng
Zhou, Daquan
Ling, Huan
Liu, Haozhe
Yi, Hongwei
Zhang, Hao
Li, Muyang
Chen, Yukang
Cai, Han
Fidler, Sanja
Luo, Ping
Han, Song
Xie, Enze
contents We introduce SANA-Video, a small diffusion model that can efficiently generate videos up to 720x1280 resolution and minute-length duration. SANA-Video synthesizes high-resolution, high-quality and long videos with strong text-video alignment at a remarkably fast speed, deployable on RTX 5090 GPU. Two core designs ensure our efficient, effective and long video generation: (1) Linear DiT: We leverage linear attention as the core operation, which is more efficient than vanilla attention given the large number of tokens processed in video generation. (2) Constant-Memory KV cache for Block Linear Attention: we design block-wise autoregressive approach for long video generation by employing a constant-memory state, derived from the cumulative properties of linear attention. This KV cache provides the Linear DiT with global context at a fixed memory cost, eliminating the need for a traditional KV cache and enabling efficient, minute-long video generation. In addition, we explore effective data filters and model training strategies, narrowing the training cost to 12 days on 64 H100 GPUs, which is only 1% of the cost of MovieGen. Given its low cost, SANA-Video achieves competitive performance compared to modern state-of-the-art small diffusion models (e.g., Wan 2.1-1.3B and SkyReel-V2-1.3B) while being 16x faster in measured latency. Moreover, SANA-Video can be deployed on RTX 5090 GPUs with NVFP4 precision, accelerating the inference speed of generating a 5-second 720p video from 71s to 29s (2.4x speedup). In summary, SANA-Video enables low-cost, high-quality video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24695
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
Chen, Junsong
Zhao, Yuyang
Yu, Jincheng
Chu, Ruihang
Chen, Junyu
Yang, Shuai
Wang, Xianbang
Pan, Yicheng
Zhou, Daquan
Ling, Huan
Liu, Haozhe
Yi, Hongwei
Zhang, Hao
Li, Muyang
Chen, Yukang
Cai, Han
Fidler, Sanja
Luo, Ping
Han, Song
Xie, Enze
Computer Vision and Pattern Recognition
Artificial Intelligence
We introduce SANA-Video, a small diffusion model that can efficiently generate videos up to 720x1280 resolution and minute-length duration. SANA-Video synthesizes high-resolution, high-quality and long videos with strong text-video alignment at a remarkably fast speed, deployable on RTX 5090 GPU. Two core designs ensure our efficient, effective and long video generation: (1) Linear DiT: We leverage linear attention as the core operation, which is more efficient than vanilla attention given the large number of tokens processed in video generation. (2) Constant-Memory KV cache for Block Linear Attention: we design block-wise autoregressive approach for long video generation by employing a constant-memory state, derived from the cumulative properties of linear attention. This KV cache provides the Linear DiT with global context at a fixed memory cost, eliminating the need for a traditional KV cache and enabling efficient, minute-long video generation. In addition, we explore effective data filters and model training strategies, narrowing the training cost to 12 days on 64 H100 GPUs, which is only 1% of the cost of MovieGen. Given its low cost, SANA-Video achieves competitive performance compared to modern state-of-the-art small diffusion models (e.g., Wan 2.1-1.3B and SkyReel-V2-1.3B) while being 16x faster in measured latency. Moreover, SANA-Video can be deployed on RTX 5090 GPUs with NVFP4 precision, accelerating the inference speed of generating a 5-second 720p video from 71s to 29s (2.4x speedup). In summary, SANA-Video enables low-cost, high-quality video generation.
title SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.24695