Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Zhiyuan, Wang, Shuai, Chen, Li, Gao, Kaihui, Li, Dan, Ren, Yanyu, Zhang, Qiming, Wang, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917132808224768
author Wu, Zhiyuan
Wang, Shuai
Chen, Li
Gao, Kaihui
Li, Dan
Ren, Yanyu
Zhang, Qiming
Wang, Yong
author_facet Wu, Zhiyuan
Wang, Shuai
Chen, Li
Gao, Kaihui
Li, Dan
Ren, Yanyu
Zhang, Qiming
Wang, Yong
contents Video diffusion models (VDMs) perform attention computation over the 3D spatio-temporal domain. Compared to large language models (LLMs) processing 1D sequences, their memory consumption scales cubically, necessitating parallel serving across multiple GPUs. Traditional parallelism strategies partition the computational graph, requiring frequent high-dimensional activation transfers that create severe communication bottlenecks. To tackle this issue, we exploit the local spatio-temporal dependencies inherent in the diffusion denoising process and propose Latent Parallelism (LP), the first parallelism strategy tailored for VDM serving. \textcolor{black}{LP decomposes the global denoising problem into parallelizable sub-problems by dynamically rotating the partitioning dimensions (temporal, height, and width) within the compact latent space across diffusion timesteps, substantially reducing the communication overhead compared to prevailing parallelism strategies.} To ensure generation quality, we design a patch-aligned overlapping partition strategy that matches partition boundaries with visual patches and a position-aware latent reconstruction mechanism for smooth stitching. Experiments on three benchmarks demonstrate that LP reduces communication overhead by up to 97\% over baseline methods while maintaining comparable generation quality. As a non-intrusive plug-in paradigm, LP can be seamlessly integrated with existing parallelism strategies, enabling efficient and scalable video generation services.
format Preprint
id arxiv_https___arxiv_org_abs_2512_07350
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism
Wu, Zhiyuan
Wang, Shuai
Chen, Li
Gao, Kaihui
Li, Dan
Ren, Yanyu
Zhang, Qiming
Wang, Yong
Distributed, Parallel, and Cluster Computing
Video diffusion models (VDMs) perform attention computation over the 3D spatio-temporal domain. Compared to large language models (LLMs) processing 1D sequences, their memory consumption scales cubically, necessitating parallel serving across multiple GPUs. Traditional parallelism strategies partition the computational graph, requiring frequent high-dimensional activation transfers that create severe communication bottlenecks. To tackle this issue, we exploit the local spatio-temporal dependencies inherent in the diffusion denoising process and propose Latent Parallelism (LP), the first parallelism strategy tailored for VDM serving. \textcolor{black}{LP decomposes the global denoising problem into parallelizable sub-problems by dynamically rotating the partitioning dimensions (temporal, height, and width) within the compact latent space across diffusion timesteps, substantially reducing the communication overhead compared to prevailing parallelism strategies.} To ensure generation quality, we design a patch-aligned overlapping partition strategy that matches partition boundaries with visual patches and a position-aware latent reconstruction mechanism for smooth stitching. Experiments on three benchmarks demonstrate that LP reduces communication overhead by up to 97\% over baseline methods while maintaining comparable generation quality. As a non-intrusive plug-in paradigm, LP can be seamlessly integrated with existing parallelism strategies, enabling efficient and scalable video generation services.
title Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2512.07350