Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Ruibin, Yang, Tao, Ai, Fangzhou, Wu, Tianhe, Wen, Shilei, Peng, Bingyue, Zhang, Lei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918470522765312
author Li, Ruibin
Yang, Tao
Ai, Fangzhou
Wu, Tianhe
Wen, Shilei
Peng, Bingyue
Zhang, Lei
author_facet Li, Ruibin
Yang, Tao
Ai, Fangzhou
Wu, Tianhe
Wen, Shilei
Peng, Bingyue
Zhang, Lei
contents Streaming video generation (SVG) distills a pretrained bidirectional video diffusion model into an autoregressive model equipped with sliding window attention (SWA). However, SWA inevitably loses distant history during long video generation, and its computational overhead remains a critical challenge to real-time deployment. In this work, we propose Hybrid Forcing, which jointly optimizes temporal information retention and computational efficiency through a hybrid attention design. First, we introduce lightweight linear temporal attention to preserve long-range dependencies beyond the sliding window. In particular, we maintain a compact key-value state to incrementally absorb evicted tokens, retaining temporal context with negligible memory and computational overhead. Second, we incorporate block-sparse attention into the local sliding window to reduce redundant computation within short-range modeling, reallocating computational capacity toward more critical dependencies. Finally, we introduce a decoupled distillation strategy tailored to the hybrid attention design. A few-step initial distillation is performed under dense attention, then the distillation of our proposed linear temporal and block-sparse attention is activated for streaming modeling, ensuring stable optimization. Extensive experiments on both short- and long-form video generation benchmarks demonstrate that Hybrid Forcing consistently achieves state-of-the-art performance. Notably, our model achieves real-time, unbounded 832x480 video generation at 29.5 FPS on a single NVIDIA H100 GPU without quantization or model compression. The source code and trained models are available at https://github.com/leeruibin/hybrid-forcing.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10103
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation
Li, Ruibin
Yang, Tao
Ai, Fangzhou
Wu, Tianhe
Wen, Shilei
Peng, Bingyue
Zhang, Lei
Computer Vision and Pattern Recognition
Streaming video generation (SVG) distills a pretrained bidirectional video diffusion model into an autoregressive model equipped with sliding window attention (SWA). However, SWA inevitably loses distant history during long video generation, and its computational overhead remains a critical challenge to real-time deployment. In this work, we propose Hybrid Forcing, which jointly optimizes temporal information retention and computational efficiency through a hybrid attention design. First, we introduce lightweight linear temporal attention to preserve long-range dependencies beyond the sliding window. In particular, we maintain a compact key-value state to incrementally absorb evicted tokens, retaining temporal context with negligible memory and computational overhead. Second, we incorporate block-sparse attention into the local sliding window to reduce redundant computation within short-range modeling, reallocating computational capacity toward more critical dependencies. Finally, we introduce a decoupled distillation strategy tailored to the hybrid attention design. A few-step initial distillation is performed under dense attention, then the distillation of our proposed linear temporal and block-sparse attention is activated for streaming modeling, ensuring stable optimization. Extensive experiments on both short- and long-form video generation benchmarks demonstrate that Hybrid Forcing consistently achieves state-of-the-art performance. Notably, our model achieves real-time, unbounded 832x480 video generation at 29.5 FPS on a single NVIDIA H100 GPU without quantization or model compression. The source code and trained models are available at https://github.com/leeruibin/hybrid-forcing.
title Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.10103