Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Ziyuan, Zheng, DanDan, Zou, Cheng, Liu, Rui, Wang, Xiaolong, Ji, Kaixiang, Chai, Weilong, Sun, Jianxin, Wang, Libin, Lv, Yongjie, Huang, Taozhi, Liu, Jiajia, Guo, Qingpei, Yang, Ming, Chen, Jingdong, Zhou, Jun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915538681200640
author Huang, Ziyuan
Zheng, DanDan
Zou, Cheng
Liu, Rui
Wang, Xiaolong
Ji, Kaixiang
Chai, Weilong
Sun, Jianxin
Wang, Libin
Lv, Yongjie
Huang, Taozhi
Liu, Jiajia
Guo, Qingpei
Yang, Ming
Chen, Jingdong
Zhou, Jun
author_facet Huang, Ziyuan
Zheng, DanDan
Zou, Cheng
Liu, Rui
Wang, Xiaolong
Ji, Kaixiang
Chai, Weilong
Sun, Jianxin
Wang, Libin
Lv, Yongjie
Huang, Taozhi
Liu, Jiajia
Guo, Qingpei
Yang, Ming
Chen, Jingdong
Zhou, Jun
contents Visual tokenization remains a core challenge in unifying visual understanding and generation within the autoregressive paradigm. Existing methods typically employ tokenizers in discrete latent spaces to align with the tokens from large language models, where the quantization errors can limit semantic expressiveness and degrade the capability of vision-language understanding. To address this, we introduce MingTok, a new family of visual tokenizers with a continuous latent space, for unified autoregressive generation and understanding. While understanding tasks favor discriminative high-dimensional features, generation tasks prefer compact low-level codes. Thus, to reconcile these competing demands, MingTok adopts a three-stage sequential architecture involving low-level encoding, semantic expansion, and visual reconstruction. Built on top of it, Ming-UniVision eliminates the need for task-specific visual representations, and unifies diverse vision-language tasks under a single autoregrsssive prediction paradigm. By formulating both understanding and generation as next-token prediction in a shared continuous space, it seamlessly supports multi-round, in-context tasks such as iterative understanding, generation and editing. Empirically, we find that using a unified continuous visual representation reconciles the competing requirements on the tokenizers by the understanding and generation tasks, thereby leading to state-of-the-art level performance across both domains. We hope our findings will facilitate unified visual tokenization in the continuous domain. Inference code and model weights are released to benefit community.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06590
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
Huang, Ziyuan
Zheng, DanDan
Zou, Cheng
Liu, Rui
Wang, Xiaolong
Ji, Kaixiang
Chai, Weilong
Sun, Jianxin
Wang, Libin
Lv, Yongjie
Huang, Taozhi
Liu, Jiajia
Guo, Qingpei
Yang, Ming
Chen, Jingdong
Zhou, Jun
Computer Vision and Pattern Recognition
Visual tokenization remains a core challenge in unifying visual understanding and generation within the autoregressive paradigm. Existing methods typically employ tokenizers in discrete latent spaces to align with the tokens from large language models, where the quantization errors can limit semantic expressiveness and degrade the capability of vision-language understanding. To address this, we introduce MingTok, a new family of visual tokenizers with a continuous latent space, for unified autoregressive generation and understanding. While understanding tasks favor discriminative high-dimensional features, generation tasks prefer compact low-level codes. Thus, to reconcile these competing demands, MingTok adopts a three-stage sequential architecture involving low-level encoding, semantic expansion, and visual reconstruction. Built on top of it, Ming-UniVision eliminates the need for task-specific visual representations, and unifies diverse vision-language tasks under a single autoregrsssive prediction paradigm. By formulating both understanding and generation as next-token prediction in a shared continuous space, it seamlessly supports multi-round, in-context tasks such as iterative understanding, generation and editing. Empirically, we find that using a unified continuous visual representation reconciles the competing requirements on the tokenizers by the understanding and generation tasks, thereby leading to state-of-the-art level performance across both domains. We hope our findings will facilitate unified visual tokenization in the continuous domain. Inference code and model weights are released to benefit community.
title Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.06590