MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Hui, Lyu, Jiayue, Wang, Fu-Yun, Cheng, Kaihui, Zhu, Siyu, Wang, Jingdong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917163668865024
author Li, Hui
Lyu, Jiayue
Wang, Fu-Yun
Cheng, Kaihui
Zhu, Siyu
Wang, Jingdong
author_facet Li, Hui
Lyu, Jiayue
Wang, Fu-Yun
Cheng, Kaihui
Zhu, Siyu
Wang, Jingdong
contents This paper studies the training-testing discrepancy (a.k.a. exposure bias) problem for improving the diffusion models. During training, the input of a prediction network at one training timestep is the corresponding ground-truth noisy data that is an interpolation of the noise and the data, and during testing, the input is the generated noisy data. We present a novel training approach, named MixFlow, for improving the performance. Our approach is motivated by the Slow Flow phenomenon: the ground-truth interpolation that is the nearest to the generated noisy data at a given sampling timestep is observed to correspond to a higher-noise timestep (termed slowed timestep), i.e., the corresponding ground-truth timestep is slower than the sampling timestep. MixFlow leverages the interpolations at the slowed timesteps, named slowed interpolation mixture, for post-training the prediction network for each training timestep. Experiments over class-conditional image generation (including SiT, REPA, and RAE) and text-to-image generation validate the effectiveness of our approach. Our approach MixFlow over the RAE models achieve strong generation results on ImageNet: 1.43 FID (without guidance) and 1.10 (with guidance) at 256 x 256, and 1.55 FID (without guidance) and 1.10 (with guidance) at 512 x 512.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19311
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
Li, Hui
Lyu, Jiayue
Wang, Fu-Yun
Cheng, Kaihui
Zhu, Siyu
Wang, Jingdong
Computer Vision and Pattern Recognition
Artificial Intelligence
This paper studies the training-testing discrepancy (a.k.a. exposure bias) problem for improving the diffusion models. During training, the input of a prediction network at one training timestep is the corresponding ground-truth noisy data that is an interpolation of the noise and the data, and during testing, the input is the generated noisy data. We present a novel training approach, named MixFlow, for improving the performance. Our approach is motivated by the Slow Flow phenomenon: the ground-truth interpolation that is the nearest to the generated noisy data at a given sampling timestep is observed to correspond to a higher-noise timestep (termed slowed timestep), i.e., the corresponding ground-truth timestep is slower than the sampling timestep. MixFlow leverages the interpolations at the slowed timesteps, named slowed interpolation mixture, for post-training the prediction network for each training timestep. Experiments over class-conditional image generation (including SiT, REPA, and RAE) and text-to-image generation validate the effectiveness of our approach. Our approach MixFlow over the RAE models achieve strong generation results on ImageNet: 1.43 FID (without guidance) and 1.10 (with guidance) at 256 x 256, and 1.55 FID (without guidance) and 1.10 (with guidance) at 512 x 512.
title MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.19311