X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Geng, Zigang, Wang, Yibing, Ma, Yeyao, Li, Chen, Rao, Yongming, Gu, Shuyang, Zhong, Zhao, Lu, Qinglin, Hu, Han, Zhang, Xiaosong, Linus, Wang, Di, Jiang, Jie
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911082041311232
author Geng, Zigang
Wang, Yibing
Ma, Yeyao
Li, Chen
Rao, Yongming
Gu, Shuyang
Zhong, Zhao
Lu, Qinglin
Hu, Han
Zhang, Xiaosong
Linus
Wang, Di
Jiang, Jie
author_facet Geng, Zigang
Wang, Yibing
Ma, Yeyao
Li, Chen
Rao, Yongming
Gu, Shuyang
Zhong, Zhao
Lu, Qinglin
Hu, Han
Zhang, Xiaosong
Linus
Wang, Di
Jiang, Jie
contents Numerous efforts have been made to extend the ``next token prediction'' paradigm to visual contents, aiming to create a unified approach for both image generation and understanding. Nevertheless, attempts to generate images through autoregressive modeling with discrete tokens have been plagued by issues such as low visual fidelity, distorted outputs, and failure to adhere to complex instructions when rendering intricate details. These shortcomings are likely attributed to cumulative errors during autoregressive inference or information loss incurred during the discretization process. Probably due to this challenge, recent research has increasingly shifted toward jointly training image generation with diffusion objectives and language generation with autoregressive objectives, moving away from unified modeling approaches. In this work, we demonstrate that reinforcement learning can effectively mitigate artifacts and largely enhance the generation quality of a discrete autoregressive modeling method, thereby enabling seamless integration of image and language generation. Our framework comprises a semantic image tokenizer, a unified autoregressive model for both language and images, and an offline diffusion decoder for image generation, termed X-Omni. X-Omni achieves state-of-the-art performance in image generation tasks using a 7B language model, producing images with high aesthetic quality while exhibiting strong capabilities in following instructions and rendering long texts.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22058
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
Geng, Zigang
Wang, Yibing
Ma, Yeyao
Li, Chen
Rao, Yongming
Gu, Shuyang
Zhong, Zhao
Lu, Qinglin
Hu, Han
Zhang, Xiaosong
Linus
Wang, Di
Jiang, Jie
Computer Vision and Pattern Recognition
Numerous efforts have been made to extend the ``next token prediction'' paradigm to visual contents, aiming to create a unified approach for both image generation and understanding. Nevertheless, attempts to generate images through autoregressive modeling with discrete tokens have been plagued by issues such as low visual fidelity, distorted outputs, and failure to adhere to complex instructions when rendering intricate details. These shortcomings are likely attributed to cumulative errors during autoregressive inference or information loss incurred during the discretization process. Probably due to this challenge, recent research has increasingly shifted toward jointly training image generation with diffusion objectives and language generation with autoregressive objectives, moving away from unified modeling approaches. In this work, we demonstrate that reinforcement learning can effectively mitigate artifacts and largely enhance the generation quality of a discrete autoregressive modeling method, thereby enabling seamless integration of image and language generation. Our framework comprises a semantic image tokenizer, a unified autoregressive model for both language and images, and an offline diffusion decoder for image generation, termed X-Omni. X-Omni achieves state-of-the-art performance in image generation tasks using a 7B language model, producing images with high aesthetic quality while exhibiting strong capabilities in following instructions and rendering long texts.
title X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.22058