One-step Latent-free Image Generation with Pixel Mean Flows

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Yiyang, Lu, Susie, Sun, Qiao, Zhao, Hanhong, Jiang, Zhicheng, Wang, Xianbang, Li, Tianhong, Geng, Zhengyang, He, Kaiming
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917476239933440
author Lu, Yiyang
Lu, Susie
Sun, Qiao
Zhao, Hanhong
Jiang, Zhicheng
Wang, Xianbang
Li, Tianhong
Geng, Zhengyang
He, Kaiming
author_facet Lu, Yiyang
Lu, Susie
Sun, Qiao
Zhao, Hanhong
Jiang, Zhicheng
Wang, Xianbang
Li, Tianhong
Geng, Zhengyang
He, Kaiming
contents Modern diffusion/flow-based models for image generation typically exhibit two core characteristics: (i) using multi-step sampling, and (ii) operating in a latent space. Recent advances have made encouraging progress on each aspect individually, paving the way toward one-step diffusion/flow without latents. In this work, we take a further step towards this goal and propose "pixel MeanFlow" (pMF). Our core guideline is to formulate the network output space and the loss space separately. The network target is designed to be on a presumed low-dimensional image manifold (i.e., x-prediction), while the loss is defined via MeanFlow in the velocity space. We introduce a simple transformation between the image manifold and the average velocity field. In experiments, pMF achieves strong results for one-step latent-free generation on ImageNet at 256x256 resolution (2.22 FID) and 512x512 resolution (2.48 FID), filling a key missing piece in this regime. We hope that our study will further advance the boundaries of diffusion/flow-based generative models.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22158
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle One-step Latent-free Image Generation with Pixel Mean Flows
Lu, Yiyang
Lu, Susie
Sun, Qiao
Zhao, Hanhong
Jiang, Zhicheng
Wang, Xianbang
Li, Tianhong
Geng, Zhengyang
He, Kaiming
Computer Vision and Pattern Recognition
Modern diffusion/flow-based models for image generation typically exhibit two core characteristics: (i) using multi-step sampling, and (ii) operating in a latent space. Recent advances have made encouraging progress on each aspect individually, paving the way toward one-step diffusion/flow without latents. In this work, we take a further step towards this goal and propose "pixel MeanFlow" (pMF). Our core guideline is to formulate the network output space and the loss space separately. The network target is designed to be on a presumed low-dimensional image manifold (i.e., x-prediction), while the loss is defined via MeanFlow in the velocity space. We introduce a simple transformation between the image manifold and the average velocity field. In experiments, pMF achieves strong results for one-step latent-free generation on ImageNet at 256x256 resolution (2.22 FID) and 512x512 resolution (2.48 FID), filling a key missing piece in this regime. We hope that our study will further advance the boundaries of diffusion/flow-based generative models.
title One-step Latent-free Image Generation with Pixel Mean Flows
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.22158