Distilling Parallel Gradients for Fast ODE Solvers of Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Beier, Wang, Ruoyu, Zhao, Tong, Zhang, Hanwang, Zhang, Chi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908457100115968
author Zhu, Beier
Wang, Ruoyu
Zhao, Tong
Zhang, Hanwang
Zhang, Chi
author_facet Zhu, Beier
Wang, Ruoyu
Zhao, Tong
Zhang, Hanwang
Zhang, Chi
contents Diffusion models (DMs) have achieved state-of-the-art generative performance but suffer from high sampling latency due to their sequential denoising nature. Existing solver-based acceleration methods often face image quality degradation under a low-latency budget. In this paper, we propose the Ensemble Parallel Direction solver (dubbed as \ours), a novel ODE solver that mitigates truncation errors by incorporating multiple parallel gradient evaluations in each ODE step. Importantly, since the additional gradient computations are independent, they can be fully parallelized, preserving low-latency sampling. Our method optimizes a small set of learnable parameters in a distillation fashion, ensuring minimal training overhead. In addition, our method can serve as a plugin to improve existing ODE samplers. Extensive experiments on various image synthesis benchmarks demonstrate the effectiveness of our \ours~in achieving high-quality and low-latency sampling. For example, at the same latency level of 5 NFE, EPD achieves an FID of 4.47 on CIFAR-10, 7.97 on FFHQ, 8.17 on ImageNet, and 8.26 on LSUN Bedroom, surpassing existing learning-based solvers by a significant margin. Codes are available in https://github.com/BeierZhu/EPD.
format Preprint
id arxiv_https___arxiv_org_abs_2507_14797
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Distilling Parallel Gradients for Fast ODE Solvers of Diffusion Models
Zhu, Beier
Wang, Ruoyu
Zhao, Tong
Zhang, Hanwang
Zhang, Chi
Computer Vision and Pattern Recognition
Diffusion models (DMs) have achieved state-of-the-art generative performance but suffer from high sampling latency due to their sequential denoising nature. Existing solver-based acceleration methods often face image quality degradation under a low-latency budget. In this paper, we propose the Ensemble Parallel Direction solver (dubbed as \ours), a novel ODE solver that mitigates truncation errors by incorporating multiple parallel gradient evaluations in each ODE step. Importantly, since the additional gradient computations are independent, they can be fully parallelized, preserving low-latency sampling. Our method optimizes a small set of learnable parameters in a distillation fashion, ensuring minimal training overhead. In addition, our method can serve as a plugin to improve existing ODE samplers. Extensive experiments on various image synthesis benchmarks demonstrate the effectiveness of our \ours~in achieving high-quality and low-latency sampling. For example, at the same latency level of 5 NFE, EPD achieves an FID of 4.47 on CIFAR-10, 7.97 on FFHQ, 8.17 on ImageNet, and 8.26 on LSUN Bedroom, surpassing existing learning-based solvers by a significant margin. Codes are available in https://github.com/BeierZhu/EPD.
title Distilling Parallel Gradients for Fast ODE Solvers of Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.14797