One-step Latent-free Image Generation with Pixel Mean Flows
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917476239933440 |
|---|---|
| author | Lu, Yiyang Lu, Susie Sun, Qiao Zhao, Hanhong Jiang, Zhicheng Wang, Xianbang Li, Tianhong Geng, Zhengyang He, Kaiming |
| author_facet | Lu, Yiyang Lu, Susie Sun, Qiao Zhao, Hanhong Jiang, Zhicheng Wang, Xianbang Li, Tianhong Geng, Zhengyang He, Kaiming |
| contents | Modern diffusion/flow-based models for image generation typically exhibit two core characteristics: (i) using multi-step sampling, and (ii) operating in a latent space. Recent advances have made encouraging progress on each aspect individually, paving the way toward one-step diffusion/flow without latents. In this work, we take a further step towards this goal and propose "pixel MeanFlow" (pMF). Our core guideline is to formulate the network output space and the loss space separately. The network target is designed to be on a presumed low-dimensional image manifold (i.e., x-prediction), while the loss is defined via MeanFlow in the velocity space. We introduce a simple transformation between the image manifold and the average velocity field. In experiments, pMF achieves strong results for one-step latent-free generation on ImageNet at 256x256 resolution (2.22 FID) and 512x512 resolution (2.48 FID), filling a key missing piece in this regime. We hope that our study will further advance the boundaries of diffusion/flow-based generative models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_22158 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | One-step Latent-free Image Generation with Pixel Mean Flows Lu, Yiyang Lu, Susie Sun, Qiao Zhao, Hanhong Jiang, Zhicheng Wang, Xianbang Li, Tianhong Geng, Zhengyang He, Kaiming Computer Vision and Pattern Recognition Modern diffusion/flow-based models for image generation typically exhibit two core characteristics: (i) using multi-step sampling, and (ii) operating in a latent space. Recent advances have made encouraging progress on each aspect individually, paving the way toward one-step diffusion/flow without latents. In this work, we take a further step towards this goal and propose "pixel MeanFlow" (pMF). Our core guideline is to formulate the network output space and the loss space separately. The network target is designed to be on a presumed low-dimensional image manifold (i.e., x-prediction), while the loss is defined via MeanFlow in the velocity space. We introduce a simple transformation between the image manifold and the average velocity field. In experiments, pMF achieves strong results for one-step latent-free generation on ImageNet at 256x256 resolution (2.22 FID) and 512x512 resolution (2.48 FID), filling a key missing piece in this regime. We hope that our study will further advance the boundaries of diffusion/flow-based generative models. |
| title | One-step Latent-free Image Generation with Pixel Mean Flows |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2601.22158 |