Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915501405372416 |
|---|---|
| author | Yue, Xiaoyu Wang, Zidong Wang, Yuqing Zhang, Wenlong Liu, Xihui Ouyang, Wanli Bai, Lei Zhou, Luping |
| author_facet | Yue, Xiaoyu Wang, Zidong Wang, Yuqing Zhang, Wenlong Liu, Xihui Ouyang, Wanli Bai, Lei Zhou, Luping |
| contents | Recent studies have demonstrated the importance of high-quality visual representations in image generation and have highlighted the limitations of generative models in image understanding. As a generative paradigm originally designed for natural language, autoregressive models face similar challenges. In this work, we present the first systematic investigation into the mechanisms of applying the next-token prediction paradigm to the visual domain. We identify three key properties that hinder the learning of high-level visual semantics: local and conditional dependence, inter-step semantic inconsistency, and spatial invariance deficiency. We show that these issues can be effectively addressed by introducing self-supervised objectives during training, leading to a novel training framework, Self-guided Training for AutoRegressive models (ST-AR). Without relying on pre-trained representation models, ST-AR significantly enhances the image understanding ability of autoregressive models and leads to improved generation quality. Specifically, ST-AR brings approximately 42% FID improvement for LlamaGen-L and 49% FID improvement for LlamaGen-XL, while maintaining the same sampling strategy. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_15185 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation Yue, Xiaoyu Wang, Zidong Wang, Yuqing Zhang, Wenlong Liu, Xihui Ouyang, Wanli Bai, Lei Zhou, Luping Computer Vision and Pattern Recognition Recent studies have demonstrated the importance of high-quality visual representations in image generation and have highlighted the limitations of generative models in image understanding. As a generative paradigm originally designed for natural language, autoregressive models face similar challenges. In this work, we present the first systematic investigation into the mechanisms of applying the next-token prediction paradigm to the visual domain. We identify three key properties that hinder the learning of high-level visual semantics: local and conditional dependence, inter-step semantic inconsistency, and spatial invariance deficiency. We show that these issues can be effectively addressed by introducing self-supervised objectives during training, leading to a novel training framework, Self-guided Training for AutoRegressive models (ST-AR). Without relying on pre-trained representation models, ST-AR significantly enhances the image understanding ability of autoregressive models and leads to improved generation quality. Specifically, ST-AR brings approximately 42% FID improvement for LlamaGen-L and 49% FID improvement for LlamaGen-XL, while maintaining the same sampling strategy. |
| title | Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.15185 |