Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Han, Jiaming, Chen, Hao, Zhao, Yang, Wang, Hanyu, Zhao, Qi, Yang, Ziyan, He, Hao, Yue, Xiangyu, Jiang, Lu
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916807662632960
author Han, Jiaming
Chen, Hao
Zhao, Yang
Wang, Hanyu
Zhao, Qi
Yang, Ziyan
He, Hao
Yue, Xiangyu
Jiang, Lu
author_facet Han, Jiaming
Chen, Hao
Zhao, Yang
Wang, Hanyu
Zhao, Qi
Yang, Ziyan
He, Hao
Yue, Xiangyu
Jiang, Lu
contents This paper presents a multimodal framework that attempts to unify visual understanding and generation within a shared discrete semantic representation. At its core is the Text-Aligned Tokenizer (TA-Tok), which converts images into discrete tokens using a text-aligned codebook projected from a large language model's (LLM) vocabulary. By integrating vision and text into a unified space with an expanded vocabulary, our multimodal LLM, Tar, enables cross-modal input and output through a shared interface, without the need for modality-specific designs. Additionally, we propose scale-adaptive encoding and decoding to balance efficiency and visual detail, along with a generative de-tokenizer to produce high-fidelity visual outputs. To address diverse decoding needs, we utilize two complementary de-tokenizers: a fast autoregressive model and a diffusion-based model. To enhance modality fusion, we investigate advanced pre-training tasks, demonstrating improvements in both visual understanding and generation. Experiments across benchmarks show that Tar matches or surpasses existing multimodal LLM methods, achieving faster convergence and greater training efficiency. Code, models, and data are available at https://tar.csuhan.com
format Preprint
id arxiv_https___arxiv_org_abs_2506_18898
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Han, Jiaming
Chen, Hao
Zhao, Yang
Wang, Hanyu
Zhao, Qi
Yang, Ziyan
He, Hao
Yue, Xiangyu
Jiang, Lu
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
This paper presents a multimodal framework that attempts to unify visual understanding and generation within a shared discrete semantic representation. At its core is the Text-Aligned Tokenizer (TA-Tok), which converts images into discrete tokens using a text-aligned codebook projected from a large language model's (LLM) vocabulary. By integrating vision and text into a unified space with an expanded vocabulary, our multimodal LLM, Tar, enables cross-modal input and output through a shared interface, without the need for modality-specific designs. Additionally, we propose scale-adaptive encoding and decoding to balance efficiency and visual detail, along with a generative de-tokenizer to produce high-fidelity visual outputs. To address diverse decoding needs, we utilize two complementary de-tokenizers: a fast autoregressive model and a diffusion-based model. To enhance modality fusion, we investigate advanced pre-training tasks, demonstrating improvements in both visual understanding and generation. Experiments across benchmarks show that Tar matches or surpasses existing multimodal LLM methods, achieving faster convergence and greater training efficiency. Code, models, and data are available at https://tar.csuhan.com
title Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
url https://arxiv.org/abs/2506.18898