Setting-Matched and Semantics-Scaled Benchmarking of One-Step Generative Models Against Multistep Diffusion and Flow Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ravishankar, Advaith, Liu, Serena, Wang, Mingyang, Zhou, Todd, Zhou, Jeffrey, Sharma, Arnav, Hu, Ziling, Das, Léopold, Sobirov, Abdulaziz, Siddique, Faizaan, Yu, Freddy, Baek, Seungjoo, Luo, Yan, Wang, Mengyu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917470639489024
author Ravishankar, Advaith
Liu, Serena
Wang, Mingyang
Zhou, Todd
Zhou, Jeffrey
Sharma, Arnav
Hu, Ziling
Das, Léopold
Sobirov, Abdulaziz
Siddique, Faizaan
Yu, Freddy
Baek, Seungjoo
Luo, Yan
Wang, Mengyu
author_facet Ravishankar, Advaith
Liu, Serena
Wang, Mingyang
Zhou, Todd
Zhou, Jeffrey
Sharma, Arnav
Hu, Ziling
Das, Léopold
Sobirov, Abdulaziz
Siddique, Faizaan
Yu, Freddy
Baek, Seungjoo
Luo, Yan
Wang, Mengyu
contents State-of-the-art text-to-image models produce high-quality images, but inference remains expensive as generation requires several sequential ODE or denoising steps. Native one-step models aim to reduce this cost by mapping noise to an image in a single step, yet fair comparisons to multi-step systems are difficult because studies use mismatched sampling steps and different classifier-free guidance (CFG) settings, where CFG can shift FID, Inception Score, and CLIP-based alignment in opposing directions. It is also unclear how well one-step models scale to multi-step inference, and there is limited standardized out-of-distribution evaluation for label-ID-conditioned generators beyond ImageNet. To address this, we benchmark eight models spanning one-step flows (MeanFlow, Improved MeanFlow, SoFlow), multi-step baselines (RAE, Scale-RAE), and established systems (SiT, Stable Diffusion 3.5, FLUX.1) under a class-conditional protocol on ImageNet validation, ImageNetV2, and reLAIONet, our new proofread out-of-distribution dataset aligned to ImageNet label IDs. Using FID, Inception Score, CLIP Score, and Pick Score, we show that FID-focused model development and CFG selection can be misleading in few-step regimes, where guidance changes can improve FID while degrading text-image alignment and human preference signals, worsening visual quality. To make these tradeoffs explicit, we introduce CLIP-scaled and PickScore-scaled variants of FID (csFID, psFID) and Inception Score (csIS, psIS) to serve as a diagnostic for semantically aligned image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_14186
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Setting-Matched and Semantics-Scaled Benchmarking of One-Step Generative Models Against Multistep Diffusion and Flow Models
Ravishankar, Advaith
Liu, Serena
Wang, Mingyang
Zhou, Todd
Zhou, Jeffrey
Sharma, Arnav
Hu, Ziling
Das, Léopold
Sobirov, Abdulaziz
Siddique, Faizaan
Yu, Freddy
Baek, Seungjoo
Luo, Yan
Wang, Mengyu
Computer Vision and Pattern Recognition
State-of-the-art text-to-image models produce high-quality images, but inference remains expensive as generation requires several sequential ODE or denoising steps. Native one-step models aim to reduce this cost by mapping noise to an image in a single step, yet fair comparisons to multi-step systems are difficult because studies use mismatched sampling steps and different classifier-free guidance (CFG) settings, where CFG can shift FID, Inception Score, and CLIP-based alignment in opposing directions. It is also unclear how well one-step models scale to multi-step inference, and there is limited standardized out-of-distribution evaluation for label-ID-conditioned generators beyond ImageNet. To address this, we benchmark eight models spanning one-step flows (MeanFlow, Improved MeanFlow, SoFlow), multi-step baselines (RAE, Scale-RAE), and established systems (SiT, Stable Diffusion 3.5, FLUX.1) under a class-conditional protocol on ImageNet validation, ImageNetV2, and reLAIONet, our new proofread out-of-distribution dataset aligned to ImageNet label IDs. Using FID, Inception Score, CLIP Score, and Pick Score, we show that FID-focused model development and CFG selection can be misleading in few-step regimes, where guidance changes can improve FID while degrading text-image alignment and human preference signals, worsening visual quality. To make these tradeoffs explicit, we introduce CLIP-scaled and PickScore-scaled variants of FID (csFID, psFID) and Inception Score (csIS, psIS) to serve as a diagnostic for semantically aligned image generation.
title Setting-Matched and Semantics-Scaled Benchmarking of One-Step Generative Models Against Multistep Diffusion and Flow Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.14186