_version_ 1866911109009637376
author NextStep Team
Han, Chunrui
Li, Guopeng
Wu, Jingwei
Sun, Quan
Cai, Yan
Peng, Yuang
Ge, Zheng
Zhou, Deyu
Tang, Haomiao
Zhou, Hongyu
Liu, Kenkun
Huang, Ailin
Wang, Bin
Miao, Changxin
Sun, Deshan
Yu, En
Yin, Fukun
Yu, Gang
Nie, Hao
Lv, Haoran
Hu, Hanpeng
Wang, Jia
Zhou, Jian
Sun, Jianjian
Tan, Kaijun
An, Kang
Lin, Kangheng
Zhao, Liang
Chen, Mei
Xing, Peng
Wang, Rui
Liu, Shiyu
Xia, Shutao
You, Tianhao
Ji, Wei
Zeng, Xianfang
Han, Xin
Zhang, Xuelin
Wei, Yana
Xu, Yanming
Jiang, Yimin
Wang, Yingming
Zhou, Yu
Han, Yucheng
Meng, Ziyang
Jiao, Binxing
Jiang, Daxin
Zhang, Xiangyu
Zhu, Yibo
author_facet NextStep Team
Han, Chunrui
Li, Guopeng
Wu, Jingwei
Sun, Quan
Cai, Yan
Peng, Yuang
Ge, Zheng
Zhou, Deyu
Tang, Haomiao
Zhou, Hongyu
Liu, Kenkun
Huang, Ailin
Wang, Bin
Miao, Changxin
Sun, Deshan
Yu, En
Yin, Fukun
Yu, Gang
Nie, Hao
Lv, Haoran
Hu, Hanpeng
Wang, Jia
Zhou, Jian
Sun, Jianjian
Tan, Kaijun
An, Kang
Lin, Kangheng
Zhao, Liang
Chen, Mei
Xing, Peng
Wang, Rui
Liu, Shiyu
Xia, Shutao
You, Tianhao
Ji, Wei
Zeng, Xianfang
Han, Xin
Zhang, Xuelin
Wei, Yana
Xu, Yanming
Jiang, Yimin
Wang, Yingming
Zhou, Yu
Han, Yucheng
Meng, Ziyang
Jiao, Binxing
Jiang, Daxin
Zhang, Xiangyu
Zhu, Yibo
contents Prevailing autoregressive (AR) models for text-to-image generation either rely on heavy, computationally-intensive diffusion models to process continuous image tokens, or employ vector quantization (VQ) to obtain discrete tokens with quantization loss. In this paper, we push the autoregressive paradigm forward with NextStep-1, a 14B autoregressive model paired with a 157M flow matching head, training on discrete text tokens and continuous image tokens with next-token prediction objectives. NextStep-1 achieves state-of-the-art performance for autoregressive models in text-to-image generation tasks, exhibiting strong capabilities in high-fidelity image synthesis. Furthermore, our method shows strong performance in image editing, highlighting the power and versatility of our unified approach. To facilitate open research, we will release our code and models to the community.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10711
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
NextStep Team
Han, Chunrui
Li, Guopeng
Wu, Jingwei
Sun, Quan
Cai, Yan
Peng, Yuang
Ge, Zheng
Zhou, Deyu
Tang, Haomiao
Zhou, Hongyu
Liu, Kenkun
Huang, Ailin
Wang, Bin
Miao, Changxin
Sun, Deshan
Yu, En
Yin, Fukun
Yu, Gang
Nie, Hao
Lv, Haoran
Hu, Hanpeng
Wang, Jia
Zhou, Jian
Sun, Jianjian
Tan, Kaijun
An, Kang
Lin, Kangheng
Zhao, Liang
Chen, Mei
Xing, Peng
Wang, Rui
Liu, Shiyu
Xia, Shutao
You, Tianhao
Ji, Wei
Zeng, Xianfang
Han, Xin
Zhang, Xuelin
Wei, Yana
Xu, Yanming
Jiang, Yimin
Wang, Yingming
Zhou, Yu
Han, Yucheng
Meng, Ziyang
Jiao, Binxing
Jiang, Daxin
Zhang, Xiangyu
Zhu, Yibo
Computer Vision and Pattern Recognition
Prevailing autoregressive (AR) models for text-to-image generation either rely on heavy, computationally-intensive diffusion models to process continuous image tokens, or employ vector quantization (VQ) to obtain discrete tokens with quantization loss. In this paper, we push the autoregressive paradigm forward with NextStep-1, a 14B autoregressive model paired with a 157M flow matching head, training on discrete text tokens and continuous image tokens with next-token prediction objectives. NextStep-1 achieves state-of-the-art performance for autoregressive models in text-to-image generation tasks, exhibiting strong capabilities in high-fidelity image synthesis. Furthermore, our method shows strong performance in image editing, highlighting the power and versatility of our unified approach. To facilitate open research, we will release our code and models to the community.
title NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.10711