DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Song, Wei, Wang, Yuran, Song, Zijia, Li, Yadong, Zhou, Zenan, Chen, Long, Xu, Jianhua, Wang, Jiaqi, Yu, Kaicheng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908978777161728
author Song, Wei
Wang, Yuran
Song, Zijia
Li, Yadong
Zhou, Zenan
Chen, Long
Xu, Jianhua
Wang, Jiaqi
Yu, Kaicheng
author_facet Song, Wei
Wang, Yuran
Song, Zijia
Li, Yadong
Zhou, Zenan
Chen, Long
Xu, Jianhua
Wang, Jiaqi
Yu, Kaicheng
contents The differing representation spaces required for visual understanding and generation pose a challenge in unifying them within the autoregressive paradigm of large language models. A vision tokenizer trained for reconstruction excels at capturing low-level visual appearance, making it well-suited for visual generation but lacking high-level semantic representations for understanding tasks. Conversely, a vision encoder trained via contrastive learning aligns well with language but struggles to decode back into the pixel space for generation tasks. To bridge this gap, we propose DualToken, a method that unifies representations for both understanding and generation within a single tokenizer. However, directly integrating reconstruction and semantic objectives creates conflicts, leading to degraded performance in both reconstruction fidelity and semantic accuracy. Instead of forcing a single codebook to capture both visual appearance and semantics, DualToken disentangles them by introducing separate codebooks for high-level semantics and low-level visual details. As a result, DualToken achieves 0.25 rFID and 82.0% zero-shot accuracy on ImageNet, and demonstrates strong effectiveness in downstream MLLM tasks for both understanding and generation. Specifically, our method surpasses VILA-U by 5.8 points on average across ten visual understanding benchmarks and delivers a 13% improvement on GenAI-Bench. Notably, incorporating dual visual tokens outperforms using a single token type on both understanding and generation tasks. We hope our research offers a new perspective on leveraging dual visual vocabularies for building unified vision-language models. Project page is available at https://songweii.github.io/dualtoken-project-page.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14324
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
Song, Wei
Wang, Yuran
Song, Zijia
Li, Yadong
Zhou, Zenan
Chen, Long
Xu, Jianhua
Wang, Jiaqi
Yu, Kaicheng
Computer Vision and Pattern Recognition
Computation and Language
The differing representation spaces required for visual understanding and generation pose a challenge in unifying them within the autoregressive paradigm of large language models. A vision tokenizer trained for reconstruction excels at capturing low-level visual appearance, making it well-suited for visual generation but lacking high-level semantic representations for understanding tasks. Conversely, a vision encoder trained via contrastive learning aligns well with language but struggles to decode back into the pixel space for generation tasks. To bridge this gap, we propose DualToken, a method that unifies representations for both understanding and generation within a single tokenizer. However, directly integrating reconstruction and semantic objectives creates conflicts, leading to degraded performance in both reconstruction fidelity and semantic accuracy. Instead of forcing a single codebook to capture both visual appearance and semantics, DualToken disentangles them by introducing separate codebooks for high-level semantics and low-level visual details. As a result, DualToken achieves 0.25 rFID and 82.0% zero-shot accuracy on ImageNet, and demonstrates strong effectiveness in downstream MLLM tasks for both understanding and generation. Specifically, our method surpasses VILA-U by 5.8 points on average across ten visual understanding benchmarks and delivers a 13% improvement on GenAI-Bench. Notably, incorporating dual visual tokens outperforms using a single token type on both understanding and generation tasks. We hope our research offers a new perspective on leveraging dual visual vocabularies for building unified vision-language models. Project page is available at https://songweii.github.io/dualtoken-project-page.
title DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2503.14324