Saved in:
Bibliographic Details
Main Authors: Zhao, Zelin, Molodyk, Petr, Xue, Haotian, Chen, Yongxin
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.19461
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912919872077824
author Zhao, Zelin
Molodyk, Petr
Xue, Haotian
Chen, Yongxin
author_facet Zhao, Zelin
Molodyk, Petr
Xue, Haotian
Chen, Yongxin
contents In this paper, we present Laplacian multiscale flow matching (LapFlow), a novel framework that enhances flow matching by leveraging multi-scale representations for image generative modeling. Our approach decomposes images into Laplacian pyramid residuals and processes different scales in parallel through a mixture-of-transformers (MoT) architecture with causal attention mechanisms. Unlike previous cascaded approaches that require explicit renoising between scales, our model generates multi-scale representations in parallel, eliminating the need for bridging processes. The proposed multi-scale architecture not only improves generation quality but also accelerates the sampling process and promotes scaling flow matching methods. Through extensive experimentation on CelebA-HQ and ImageNet, we demonstrate that our method achieves superior sample quality with fewer GFLOPs and faster inference compared to single-scale and multi-scale flow matching baselines. The proposed model scales effectively to high-resolution generation (up to 1024$\times$1024) while maintaining lower computational overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2602_19461
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Laplacian Multi-scale Flow Matching for Generative Modeling
Zhao, Zelin
Molodyk, Petr
Xue, Haotian
Chen, Yongxin
Computer Vision and Pattern Recognition
Machine Learning
In this paper, we present Laplacian multiscale flow matching (LapFlow), a novel framework that enhances flow matching by leveraging multi-scale representations for image generative modeling. Our approach decomposes images into Laplacian pyramid residuals and processes different scales in parallel through a mixture-of-transformers (MoT) architecture with causal attention mechanisms. Unlike previous cascaded approaches that require explicit renoising between scales, our model generates multi-scale representations in parallel, eliminating the need for bridging processes. The proposed multi-scale architecture not only improves generation quality but also accelerates the sampling process and promotes scaling flow matching methods. Through extensive experimentation on CelebA-HQ and ImageNet, we demonstrate that our method achieves superior sample quality with fewer GFLOPs and faster inference compared to single-scale and multi-scale flow matching baselines. The proposed model scales effectively to high-resolution generation (up to 1024$\times$1024) while maintaining lower computational overhead.
title Laplacian Multi-scale Flow Matching for Generative Modeling
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2602.19461