TReFT: Taming Rectified Flow Models For One-Step Image Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Shengqian, Gao, Ming, Liu, Yi, Lin, Zuzeng, Wang, Feng, Dai, Feng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911286199058432
author Li, Shengqian
Gao, Ming
Liu, Yi
Lin, Zuzeng
Wang, Feng
Dai, Feng
author_facet Li, Shengqian
Gao, Ming
Liu, Yi
Lin, Zuzeng
Wang, Feng
Dai, Feng
contents Rectified Flow (RF) models have advanced high-quality image and video synthesis via optimal transport theory. However, when applied to image-to-image translation, they still depend on costly multi-step denoising, hindering real-time applications. Although the recent adversarial training paradigm, CycleGAN-Turbo, works in pretrained diffusion models for one-step image translation, we find that directly applying it to RF models leads to severe convergence issues. In this paper, we analyze these challenges and propose TReFT, a novel method to Tame Rectified Flow models for one-step image Translation. Unlike previous works, TReFT directly uses the velocity predicted by pretrained DiT or UNet as output-a simple yet effective design that tackles the convergence issues under adversarial training with one-step inference. This design is mainly motivated by a novel observation that, near the end of the denoising process, the velocity predicted by pretrained RF models converges to the vector from origin to the final clean image, a property we further justify through theoretical analysis. When applying TReFT to large pretrained RF models such as SD3.5 and FLUX, we introduce memory-efficient latent cycle-consistency and identity losses during training, as well as lightweight architectural simplifications for faster inference. Pretrained RF models finetuned with TReFT achieve performance comparable to sota methods across multiple image translation datasets while enabling real-time inference.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20307
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TReFT: Taming Rectified Flow Models For One-Step Image Translation
Li, Shengqian
Gao, Ming
Liu, Yi
Lin, Zuzeng
Wang, Feng
Dai, Feng
Computer Vision and Pattern Recognition
Rectified Flow (RF) models have advanced high-quality image and video synthesis via optimal transport theory. However, when applied to image-to-image translation, they still depend on costly multi-step denoising, hindering real-time applications. Although the recent adversarial training paradigm, CycleGAN-Turbo, works in pretrained diffusion models for one-step image translation, we find that directly applying it to RF models leads to severe convergence issues. In this paper, we analyze these challenges and propose TReFT, a novel method to Tame Rectified Flow models for one-step image Translation. Unlike previous works, TReFT directly uses the velocity predicted by pretrained DiT or UNet as output-a simple yet effective design that tackles the convergence issues under adversarial training with one-step inference. This design is mainly motivated by a novel observation that, near the end of the denoising process, the velocity predicted by pretrained RF models converges to the vector from origin to the final clean image, a property we further justify through theoretical analysis. When applying TReFT to large pretrained RF models such as SD3.5 and FLUX, we introduce memory-efficient latent cycle-consistency and identity losses during training, as well as lightweight architectural simplifications for faster inference. Pretrained RF models finetuned with TReFT achieve performance comparable to sota methods across multiple image translation datasets while enabling real-time inference.
title TReFT: Taming Rectified Flow Models For One-Step Image Translation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.20307