UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tang, Hongxuan, Liu, Hao, Xiao, Xinyan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910895425191936
author Tang, Hongxuan
Liu, Hao
Xiao, Xinyan
author_facet Tang, Hongxuan
Liu, Hao
Xiao, Xinyan
contents We introduce UGen, a unified autoregressive multimodal model that demonstrates strong performance across text processing, image understanding, and image generation tasks simultaneously. UGen converts both texts and images into discrete token sequences and utilizes a single transformer to generate them uniformly in an autoregressive manner. To address the challenges associated with unified multimodal learning, UGen is trained using a novel mechanism, namely progressive vocabulary learning. In this process, visual token IDs are incrementally activated and integrated into the training phase, ultimately enhancing the effectiveness of unified multimodal learning. Experiments on comprehensive text and image tasks show that UGen achieves a significant overall performance improvement of 13.3% compared to the vanilla unified autoregressive method, and it also delivers competitive results across all tasks against several task-specific models.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21193
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning
Tang, Hongxuan
Liu, Hao
Xiao, Xinyan
Computation and Language
Computer Vision and Pattern Recognition
We introduce UGen, a unified autoregressive multimodal model that demonstrates strong performance across text processing, image understanding, and image generation tasks simultaneously. UGen converts both texts and images into discrete token sequences and utilizes a single transformer to generate them uniformly in an autoregressive manner. To address the challenges associated with unified multimodal learning, UGen is trained using a novel mechanism, namely progressive vocabulary learning. In this process, visual token IDs are incrementally activated and integrated into the training phase, ultimately enhancing the effectiveness of unified multimodal learning. Experiments on comprehensive text and image tasks show that UGen achieves a significant overall performance improvement of 13.3% compared to the vanilla unified autoregressive method, and it also delivers competitive results across all tasks against several task-specific models.
title UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.21193