HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qiu, Haonan, Liu, Shikun, Zhou, Zijian, An, Zhaochong, Ren, Weiming, Liu, Zhiheng, Schult, Jonas, He, Sen, Chen, Shoufa, Cong, Yuren, Xiang, Tao, Liu, Ziwei, Perez-Rua, Juan-Manuel
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908731645624320
author Qiu, Haonan
Liu, Shikun
Zhou, Zijian
An, Zhaochong
Ren, Weiming
Liu, Zhiheng
Schult, Jonas
He, Sen
Chen, Shoufa
Cong, Yuren
Xiang, Tao
Liu, Ziwei
Perez-Rua, Juan-Manuel
author_facet Qiu, Haonan
Liu, Shikun
Zhou, Zijian
An, Zhaochong
Ren, Weiming
Liu, Zhiheng
Schult, Jonas
He, Sen
Chen, Shoufa
Cong, Yuren
Xiang, Tao
Liu, Ziwei
Perez-Rua, Juan-Manuel
contents High-resolution video generation, while crucial for digital media and film, is computationally bottlenecked by the quadratic complexity of diffusion models, making practical inference infeasible. To address this, we introduce HiStream, an efficient autoregressive framework that systematically reduces redundancy across three axes: i) Spatial Compression: denoising at low resolution before refining at high resolution with cached features; ii) Temporal Compression: a chunk-by-chunk strategy with a fixed-size anchor cache, ensuring stable inference speed; and iii) Timestep Compression: applying fewer denoising steps to subsequent, cache-conditioned chunks. On 1080p benchmarks, our primary HiStream model (i+ii) achieves state-of-the-art visual quality while demonstrating up to 76.2x faster denoising compared to the Wan2.1 baseline and negligible quality loss. Our faster variant, HiStream+, applies all three optimizations (i+ii+iii), achieving a 107.5x acceleration over the baseline, offering a compelling trade-off between speed and quality, thereby making high-resolution video generation both practical and scalable.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21338
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
Qiu, Haonan
Liu, Shikun
Zhou, Zijian
An, Zhaochong
Ren, Weiming
Liu, Zhiheng
Schult, Jonas
He, Sen
Chen, Shoufa
Cong, Yuren
Xiang, Tao
Liu, Ziwei
Perez-Rua, Juan-Manuel
Computer Vision and Pattern Recognition
High-resolution video generation, while crucial for digital media and film, is computationally bottlenecked by the quadratic complexity of diffusion models, making practical inference infeasible. To address this, we introduce HiStream, an efficient autoregressive framework that systematically reduces redundancy across three axes: i) Spatial Compression: denoising at low resolution before refining at high resolution with cached features; ii) Temporal Compression: a chunk-by-chunk strategy with a fixed-size anchor cache, ensuring stable inference speed; and iii) Timestep Compression: applying fewer denoising steps to subsequent, cache-conditioned chunks. On 1080p benchmarks, our primary HiStream model (i+ii) achieves state-of-the-art visual quality while demonstrating up to 76.2x faster denoising compared to the Wan2.1 baseline and negligible quality loss. Our faster variant, HiStream+, applies all three optimizations (i+ii+iii), achieving a 107.5x acceleration over the baseline, offering a compelling trade-off between speed and quality, thereby making high-resolution video generation both practical and scalable.
title HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.21338