Improving Progressive Generation with Decomposable Flow Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Haji-Ali, Moayed, Menapace, Willi, Skorokhodov, Ivan, Sahni, Arpit, Tulyakov, Sergey, Ordonez, Vicente, Siarohin, Aliaksandr
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913910467067904
author Haji-Ali, Moayed
Menapace, Willi
Skorokhodov, Ivan
Sahni, Arpit
Tulyakov, Sergey
Ordonez, Vicente
Siarohin, Aliaksandr
author_facet Haji-Ali, Moayed
Menapace, Willi
Skorokhodov, Ivan
Sahni, Arpit
Tulyakov, Sergey
Ordonez, Vicente
Siarohin, Aliaksandr
contents Generating high-dimensional visual modalities is a computationally intensive task. A common solution is progressive generation, where the outputs are synthesized in a coarse-to-fine spectral autoregressive manner. While diffusion models benefit from the coarse-to-fine nature of denoising, explicit multi-stage architectures are rarely adopted. These architectures have increased the complexity of the overall approach, introducing the need for a custom diffusion formulation, decomposition-dependent stage transitions, add-hoc samplers, or a model cascade. Our contribution, Decomposable Flow Matching (DFM), is a simple and effective framework for the progressive generation of visual media. DFM applies Flow Matching independently at each level of a user-defined multi-scale representation (such as Laplacian pyramid). As shown by our experiments, our approach improves visual quality for both images and videos, featuring superior results compared to prior multistage frameworks. On Imagenet-1k 512px, DFM achieves 35.2% improvements in FDD scores over the base architecture and 26.4% over the best-performing baseline, under the same training compute. When applied to finetuning of large models, such as FLUX, DFM shows faster convergence speed to the training distribution. Crucially, all these advantages are achieved with a single model, architectural simplicity, and minimal modifications to existing training pipelines.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19839
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Progressive Generation with Decomposable Flow Matching
Haji-Ali, Moayed
Menapace, Willi
Skorokhodov, Ivan
Sahni, Arpit
Tulyakov, Sergey
Ordonez, Vicente
Siarohin, Aliaksandr
Computer Vision and Pattern Recognition
Artificial Intelligence
Generating high-dimensional visual modalities is a computationally intensive task. A common solution is progressive generation, where the outputs are synthesized in a coarse-to-fine spectral autoregressive manner. While diffusion models benefit from the coarse-to-fine nature of denoising, explicit multi-stage architectures are rarely adopted. These architectures have increased the complexity of the overall approach, introducing the need for a custom diffusion formulation, decomposition-dependent stage transitions, add-hoc samplers, or a model cascade. Our contribution, Decomposable Flow Matching (DFM), is a simple and effective framework for the progressive generation of visual media. DFM applies Flow Matching independently at each level of a user-defined multi-scale representation (such as Laplacian pyramid). As shown by our experiments, our approach improves visual quality for both images and videos, featuring superior results compared to prior multistage frameworks. On Imagenet-1k 512px, DFM achieves 35.2% improvements in FDD scores over the base architecture and 26.4% over the best-performing baseline, under the same training compute. When applied to finetuning of large models, such as FLUX, DFM shows faster convergence speed to the training distribution. Crucially, all these advantages are achieved with a single model, architectural simplicity, and minimal modifications to existing training pipelines.
title Improving Progressive Generation with Decomposable Flow Matching
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.19839