UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Yanzhe, Zhong, Huasong, Li, Yan, Yang, Zhenheng
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912471275536384
author Chen, Yanzhe
Zhong, Huasong
Li, Yan
Yang, Zhenheng
author_facet Chen, Yanzhe
Zhong, Huasong
Li, Yan
Yang, Zhenheng
contents Unified multimodal large language models (MLLMs) have shown promise in jointly advancing multimodal understanding and generation, with visual codebooks discretizing images into tokens for autoregressive modeling. Existing codebook-based methods either rely on small vocabularies (~16K entries) that lack fine-grained semantics or naively scale up, resulting in low token utilization and unstable training. We propose UniCode$^2$, a cascaded codebook framework enabling large-scale, semantically aligned, and stable visual tokenization. By clustering millions of SigLIP sequence embeddings, we build a 500K-entry codebook that preserves vision-language alignment while expanding capacity. Stability is ensured via a cascaded design: a frozen codebook anchors the embedding space, and a trainable codebook refines task-specific semantics. This decoupling promotes high utilization and robust learning. Moreover, the alignment of our visual tokens with textual semantics enables seamless integration with pretrained diffusion decoders, supporting high-quality visual synthesis with minimal adaptation. UniCode^2 delivers strong performance across diverse benchmarks, demonstrating the viability of scaling visual token spaces without sacrificing stability, semantics, or modularity.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20214
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
Chen, Yanzhe
Zhong, Huasong
Li, Yan
Yang, Zhenheng
Computer Vision and Pattern Recognition
Multimedia
Unified multimodal large language models (MLLMs) have shown promise in jointly advancing multimodal understanding and generation, with visual codebooks discretizing images into tokens for autoregressive modeling. Existing codebook-based methods either rely on small vocabularies (~16K entries) that lack fine-grained semantics or naively scale up, resulting in low token utilization and unstable training. We propose UniCode$^2$, a cascaded codebook framework enabling large-scale, semantically aligned, and stable visual tokenization. By clustering millions of SigLIP sequence embeddings, we build a 500K-entry codebook that preserves vision-language alignment while expanding capacity. Stability is ensured via a cascaded design: a frozen codebook anchors the embedding space, and a trainable codebook refines task-specific semantics. This decoupling promotes high utilization and robust learning. Moreover, the alignment of our visual tokens with textual semantics enables seamless integration with pretrained diffusion decoders, supporting high-quality visual synthesis with minimal adaptation. UniCode^2 delivers strong performance across diverse benchmarks, demonstrating the viability of scaling visual token spaces without sacrificing stability, semantics, or modularity.
title UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2506.20214