DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yiheng, Qu, Liao, Zhang, Huichao, Wang, Xu, Jiang, Yi, Gao, Yiming, Ye, Hu, Li, Xian, Wang, Shuai, Du, Daniel K., Chen, Fangmin, Yuan, Zehuan, Wu, Xinglong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909897297231872
author Liu, Yiheng
Qu, Liao
Zhang, Huichao
Wang, Xu
Jiang, Yi
Gao, Yiming
Ye, Hu
Li, Xian
Wang, Shuai
Du, Daniel K.
Chen, Fangmin
Yuan, Zehuan
Wu, Xinglong
author_facet Liu, Yiheng
Qu, Liao
Zhang, Huichao
Wang, Xu
Jiang, Yi
Gao, Yiming
Ye, Hu
Li, Xian
Wang, Shuai
Du, Daniel K.
Chen, Fangmin
Yuan, Zehuan
Wu, Xinglong
contents This paper presents DetailFlow, a coarse-to-fine 1D autoregressive (AR) image generation method that models images through a novel next-detail prediction strategy. By learning a resolution-aware token sequence supervised with progressively degraded images, DetailFlow enables the generation process to start from the global structure and incrementally refine details. This coarse-to-fine 1D token sequence aligns well with the autoregressive inference mechanism, providing a more natural and efficient way for the AR model to generate complex visual content. Our compact 1D AR model achieves high-quality image synthesis with significantly fewer tokens than previous approaches, i.e. VAR/VQGAN. We further propose a parallel inference mechanism with self-correction that accelerates generation speed by approximately 8x while reducing accumulation sampling error inherent in teacher-forcing supervision. On the ImageNet 256x256 benchmark, our method achieves 2.96 gFID with 128 tokens, outperforming VAR (3.3 FID) and FlexVAR (3.05 FID), which both require 680 tokens in their AR models. Moreover, due to the significantly reduced token count and parallel inference mechanism, our method runs nearly 2x faster inference speed compared to VAR and FlexVAR. Extensive experimental results demonstrate DetailFlow's superior generation quality and efficiency compared to existing state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21473
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction
Liu, Yiheng
Qu, Liao
Zhang, Huichao
Wang, Xu
Jiang, Yi
Gao, Yiming
Ye, Hu
Li, Xian
Wang, Shuai
Du, Daniel K.
Chen, Fangmin
Yuan, Zehuan
Wu, Xinglong
Computer Vision and Pattern Recognition
This paper presents DetailFlow, a coarse-to-fine 1D autoregressive (AR) image generation method that models images through a novel next-detail prediction strategy. By learning a resolution-aware token sequence supervised with progressively degraded images, DetailFlow enables the generation process to start from the global structure and incrementally refine details. This coarse-to-fine 1D token sequence aligns well with the autoregressive inference mechanism, providing a more natural and efficient way for the AR model to generate complex visual content. Our compact 1D AR model achieves high-quality image synthesis with significantly fewer tokens than previous approaches, i.e. VAR/VQGAN. We further propose a parallel inference mechanism with self-correction that accelerates generation speed by approximately 8x while reducing accumulation sampling error inherent in teacher-forcing supervision. On the ImageNet 256x256 benchmark, our method achieves 2.96 gFID with 128 tokens, outperforming VAR (3.3 FID) and FlexVAR (3.05 FID), which both require 680 tokens in their AR models. Moreover, due to the significantly reduced token count and parallel inference mechanism, our method runs nearly 2x faster inference speed compared to VAR and FlexVAR. Extensive experimental results demonstrate DetailFlow's superior generation quality and efficiency compared to existing state-of-the-art methods.
title DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.21473