LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yukang, Wang, Luozhou, Huang, Wei, Yang, Shuai, Zhang, Bohan, Xiao, Yicheng, Chu, Ruihang, Mao, Weian, Hu, Qixin, Liu, Shaoteng, Zhao, Yuyang, Mao, Huizi, Chen, Ying-Cong, Xie, Enze, Qi, Xiaojuan, Han, Song
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914580207239168
author Chen, Yukang
Wang, Luozhou
Huang, Wei
Yang, Shuai
Zhang, Bohan
Xiao, Yicheng
Chu, Ruihang
Mao, Weian
Hu, Qixin
Liu, Shaoteng
Zhao, Yuyang
Mao, Huizi
Chen, Ying-Cong
Xie, Enze
Qi, Xiaojuan
Han, Song
author_facet Chen, Yukang
Wang, Luozhou
Huang, Wei
Yang, Shuai
Zhang, Bohan
Xiao, Yicheng
Chu, Ruihang
Mao, Weian
Hu, Qixin
Liu, Shaoteng
Zhao, Yuyang
Mao, Huizi
Chen, Ying-Cong
Xie, Enze
Qi, Xiaojuan
Han, Song
contents We present LongLive-2.0, an NVFP4-based parallel infrastructure throughout the full training and inference workflow of long video generation, addressing speed and memory bottlenecks. For training, we introduce sequence-parallel autoregressive (AR) training, instantiated as Balanced SP, which co-designs the efficient teacher-forcing layout with SP execution by pairing clean-history and noisy-target temporal chunks on each rank, enabling a natural teacher-forcing mask with SP-aware chunked VAE encoding. Combined with NVFP4 precision, it reduces GPU memory cost and accelerates GEMM computation during training, the proportion of which increases as video length grows. Moreover, we show that a high-quality infrastructure and dataset enable a remarkably clean training pipeline. Unlike existing Self-Forcing series methods that rely on ODE initialization and subsequent distribution matching distillation (DMD), LongLive-2.0 directly tunes a diffusion model into a long, multi-shot, interactive auto-regressive (AR) diffusion model. It can be further converted to real-time generation (4 to 2 denoising steps) with standalone LoRA weights. For inference on Blackwell GPUs, we enable W4A4 NVFP4 inference, quantize KV cache into NVFP4 for memory savings, and boost end-to-end throughput with asynchronous streaming VAE decoding. On non-Blackwell GPU architectures, we deploy SP inference to match the speed on Blackwell GPUs, while the quantized KV cache can lower inter-GPU communication of SP. Experiments show up to 2.15x speedup in training, and 1.84x in inference. LongLive-2.0-5B achieves 45.7 FPS inference while attaining strong performance on benchmarks. To our knowledge, LongLive-2.0 is the first NVFP4 training and inference system for long video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18739
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
Chen, Yukang
Wang, Luozhou
Huang, Wei
Yang, Shuai
Zhang, Bohan
Xiao, Yicheng
Chu, Ruihang
Mao, Weian
Hu, Qixin
Liu, Shaoteng
Zhao, Yuyang
Mao, Huizi
Chen, Ying-Cong
Xie, Enze
Qi, Xiaojuan
Han, Song
Computer Vision and Pattern Recognition
Distributed, Parallel, and Cluster Computing
We present LongLive-2.0, an NVFP4-based parallel infrastructure throughout the full training and inference workflow of long video generation, addressing speed and memory bottlenecks. For training, we introduce sequence-parallel autoregressive (AR) training, instantiated as Balanced SP, which co-designs the efficient teacher-forcing layout with SP execution by pairing clean-history and noisy-target temporal chunks on each rank, enabling a natural teacher-forcing mask with SP-aware chunked VAE encoding. Combined with NVFP4 precision, it reduces GPU memory cost and accelerates GEMM computation during training, the proportion of which increases as video length grows. Moreover, we show that a high-quality infrastructure and dataset enable a remarkably clean training pipeline. Unlike existing Self-Forcing series methods that rely on ODE initialization and subsequent distribution matching distillation (DMD), LongLive-2.0 directly tunes a diffusion model into a long, multi-shot, interactive auto-regressive (AR) diffusion model. It can be further converted to real-time generation (4 to 2 denoising steps) with standalone LoRA weights. For inference on Blackwell GPUs, we enable W4A4 NVFP4 inference, quantize KV cache into NVFP4 for memory savings, and boost end-to-end throughput with asynchronous streaming VAE decoding. On non-Blackwell GPU architectures, we deploy SP inference to match the speed on Blackwell GPUs, while the quantized KV cache can lower inter-GPU communication of SP. Experiments show up to 2.15x speedup in training, and 1.84x in inference. LongLive-2.0-5B achieves 45.7 FPS inference while attaining strong performance on benchmarks. To our knowledge, LongLive-2.0 is the first NVFP4 training and inference system for long video generation.
title LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
topic Computer Vision and Pattern Recognition
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2605.18739