Staleness-Centric Optimizations for Parallel Diffusion MoE Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Jiajun, Luo, Lizhuo, Xu, Jianru, Song, Jiajun, Lu, Rongwei, Tang, Chen, Wang, Zhi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914172792471552
author Luo, Jiajun
Luo, Lizhuo
Xu, Jianru
Song, Jiajun
Lu, Rongwei
Tang, Chen
Wang, Zhi
author_facet Luo, Jiajun
Luo, Lizhuo
Xu, Jianru
Song, Jiajun
Lu, Rongwei
Tang, Chen
Wang, Zhi
contents Mixture-of-Experts-based (MoE-based) diffusion models demonstrate remarkable scalability in high-fidelity image generation, yet their reliance on expert parallelism introduces critical communication bottlenecks. State-of-the-art methods alleviate such overhead in parallel diffusion inference through computation-communication overlapping, termed displaced parallelism. However, we identify that these techniques induce severe *staleness*-the usage of outdated activations from previous timesteps that significantly degrades quality, especially in expert-parallel scenarios. We tackle this fundamental tension and propose DICE, a staleness-centric optimization framework with a three-fold approach: (1) Interweaved Parallelism introduces staggered pipelines, effectively halving step-level staleness for free; (2) Selective Synchronization operates at layer-level and protects layers vulnerable from staled activations; and (3) Conditional Communication, a token-level, training-free method that dynamically adjusts communication frequency based on token importance. Together, these strategies effectively reduce staleness, achieving 1.26x speedup with minimal quality degradation. Empirical results establish DICE as an effective and scalable solution. Our code is publicly available at https://github.com/Cobalt-27/DICE
format Preprint
id arxiv_https___arxiv_org_abs_2411_16786
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
Luo, Jiajun
Luo, Lizhuo
Xu, Jianru
Song, Jiajun
Lu, Rongwei
Tang, Chen
Wang, Zhi
Distributed, Parallel, and Cluster Computing
Mixture-of-Experts-based (MoE-based) diffusion models demonstrate remarkable scalability in high-fidelity image generation, yet their reliance on expert parallelism introduces critical communication bottlenecks. State-of-the-art methods alleviate such overhead in parallel diffusion inference through computation-communication overlapping, termed displaced parallelism. However, we identify that these techniques induce severe *staleness*-the usage of outdated activations from previous timesteps that significantly degrades quality, especially in expert-parallel scenarios. We tackle this fundamental tension and propose DICE, a staleness-centric optimization framework with a three-fold approach: (1) Interweaved Parallelism introduces staggered pipelines, effectively halving step-level staleness for free; (2) Selective Synchronization operates at layer-level and protects layers vulnerable from staled activations; and (3) Conditional Communication, a token-level, training-free method that dynamically adjusts communication frequency based on token importance. Together, these strategies effectively reduce staleness, achieving 1.26x speedup with minimal quality degradation. Empirical results establish DICE as an effective and scalable solution. Our code is publicly available at https://github.com/Cobalt-27/DICE
title Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2411.16786