NAMI: Efficient Image Generation via Bridged Progressive Rectified Flow Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Yuhang, Cheng, Bo, Liu, Shanyuan, Zhou, Hongyi, Wu, Liebucha, Leng, Dawei, Yin, Yuhui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912946582454272
author Ma, Yuhang
Cheng, Bo
Liu, Shanyuan
Zhou, Hongyi
Wu, Liebucha
Leng, Dawei
Yin, Yuhui
author_facet Ma, Yuhang
Cheng, Bo
Liu, Shanyuan
Zhou, Hongyi
Wu, Liebucha
Leng, Dawei
Yin, Yuhui
contents Flow-based Transformer models have achieved state-of-the-art image generation performance, but often suffer from high inference latency and computational cost due to their large parameter sizes. To improve inference efficiency without compromising quality, we propose Bridged Progressive Rectified Flow Transformers (NAMI), which decompose the generation process across temporal, spatial, and architectural demensions. We divide the rectified flow into different stages according to resolution, and use a BridgeFlow module to connect them. Fewer Transformer layers are used at low-resolution stages to generate image layouts and concept contours, and more layers are progressively added as the resolution increases. Experiments demonstrate that our approach achieves fast convergence and reduces inference time while ensuring generation quality. The main contributions of this paper are summarized as follows: (1) We introduce Bridged Progressive Rectified Flow Transformers that enable multi-resolution training, accelerating model convergence; (2) NAMI leverages piecewise flow and spatial cascading of Diffusion Transformer (DiT) to rapidly generate images, reducing inference time by 64% for generating 1024 resolution images; (3) We propose a BridgeFlow module to align flows between different stages; (4) We propose the NAMI-1K benchmark to evaluate human preference performance, aiming to mitigate distributional bias and comprehensively assess model effectiveness. The results show that our model is competitive with state-of-the-art models.
format Preprint
id arxiv_https___arxiv_org_abs_2503_09242
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NAMI: Efficient Image Generation via Bridged Progressive Rectified Flow Transformers
Ma, Yuhang
Cheng, Bo
Liu, Shanyuan
Zhou, Hongyi
Wu, Liebucha
Leng, Dawei
Yin, Yuhui
Computer Vision and Pattern Recognition
Flow-based Transformer models have achieved state-of-the-art image generation performance, but often suffer from high inference latency and computational cost due to their large parameter sizes. To improve inference efficiency without compromising quality, we propose Bridged Progressive Rectified Flow Transformers (NAMI), which decompose the generation process across temporal, spatial, and architectural demensions. We divide the rectified flow into different stages according to resolution, and use a BridgeFlow module to connect them. Fewer Transformer layers are used at low-resolution stages to generate image layouts and concept contours, and more layers are progressively added as the resolution increases. Experiments demonstrate that our approach achieves fast convergence and reduces inference time while ensuring generation quality. The main contributions of this paper are summarized as follows: (1) We introduce Bridged Progressive Rectified Flow Transformers that enable multi-resolution training, accelerating model convergence; (2) NAMI leverages piecewise flow and spatial cascading of Diffusion Transformer (DiT) to rapidly generate images, reducing inference time by 64% for generating 1024 resolution images; (3) We propose a BridgeFlow module to align flows between different stages; (4) We propose the NAMI-1K benchmark to evaluate human preference performance, aiming to mitigate distributional bias and comprehensively assess model effectiveness. The results show that our model is competitive with state-of-the-art models.
title NAMI: Efficient Image Generation via Bridged Progressive Rectified Flow Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.09242