VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Junwen, Liang, Jinglin, Chen, Hongyuan, Huang, Shuangping
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913155295215616
author Tan, Junwen
Liang, Jinglin
Chen, Hongyuan
Huang, Shuangping
author_facet Tan, Junwen
Liang, Jinglin
Chen, Hongyuan
Huang, Shuangping
contents Though rectified flow models have achieved remarkable performance in image, video, and 3D generation, their practical deployments are challenged by slow inference speeds. Prior acceleration methods reuse cached features from previous steps, which neglects the growing mismatch between static caches and the evolving input, leading to reduced output fidelity. This work proposes Velocity Decomposition and Estimation (VDE), a training-free acceleration method that shifts the paradigm from caching-and-reusing to decomposing-and-estimating. Specifically, VDE decomposes the model's velocity into components parallel and orthogonal to the input, exploiting their temporal predictability and directional stability for precise, input-adaptive estimation. To prevent error accumulation, it periodically anchors the model's state via full forward passes. Extensive experiments on image and video generation tasks demonstrate that VDE achieves substantial acceleration with minimal loss in visual quality. Notably, VDE accelerates Flux by 3.22 times and achieves an LPIPS of 0.069 on Qwen-Image, outperforming the best baseline with a 52.2% reduction.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23381
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and Estimation
Tan, Junwen
Liang, Jinglin
Chen, Hongyuan
Huang, Shuangping
Computer Vision and Pattern Recognition
Though rectified flow models have achieved remarkable performance in image, video, and 3D generation, their practical deployments are challenged by slow inference speeds. Prior acceleration methods reuse cached features from previous steps, which neglects the growing mismatch between static caches and the evolving input, leading to reduced output fidelity. This work proposes Velocity Decomposition and Estimation (VDE), a training-free acceleration method that shifts the paradigm from caching-and-reusing to decomposing-and-estimating. Specifically, VDE decomposes the model's velocity into components parallel and orthogonal to the input, exploiting their temporal predictability and directional stability for precise, input-adaptive estimation. To prevent error accumulation, it periodically anchors the model's state via full forward passes. Extensive experiments on image and video generation tasks demonstrate that VDE achieves substantial acceleration with minimal loss in visual quality. Notably, VDE accelerates Flux by 3.22 times and achieves an LPIPS of 0.069 on Qwen-Image, outperforming the best baseline with a 52.2% reduction.
title VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.23381