A More Word-like Image Tokenization for MLLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lee, Hyun, Jeong, Hyemin, Kim, Yejin, Choi, Hyungwook, Cho, Hyunsoo, Kim, Soo Kyung, Lee, Joonseok
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914577366646784
author Lee, Hyun
Jeong, Hyemin
Kim, Yejin
Choi, Hyungwook
Cho, Hyunsoo
Kim, Soo Kyung
Lee, Joonseok
author_facet Lee, Hyun
Jeong, Hyemin
Kim, Yejin
Choi, Hyungwook
Cho, Hyunsoo
Kim, Soo Kyung
Lee, Joonseok
contents Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a sequence of tokens in its embedding space, so that images can be presented in essentially the same form as text. However, the language model has been optimized to operate on discrete, semantically meaningful tokens, while prevailing visual projectors transform an image into a long stream of continuous and highly correlated embeddings. This causes the visual tokens to behave differently from the word-like units that LLMs are originally trained to understand. We propose a novel Disentangled Visual Tokenization (DiVT) that clusters patch embeddings into coherent semantic units, so each token corresponds to a distinct visual concept instead of a rigid grid cell. DiVT further adapts its token budget to image complexity, providing an explicit accuracy-compute trade-off modifying neither the vision encoder nor the language model. Across diverse multimodal benchmarks, DiVT matches or surpasses baselines with significantly fewer visual tokens, demonstrating robustness under limited token budgets, significantly reducing memory cost and latency while making visual inputs more compatible with LLMs. Our code is available at https://github.com/snuviplab/DiVT.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17954
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A More Word-like Image Tokenization for MLLMs
Lee, Hyun
Jeong, Hyemin
Kim, Yejin
Choi, Hyungwook
Cho, Hyunsoo
Kim, Soo Kyung
Lee, Joonseok
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a sequence of tokens in its embedding space, so that images can be presented in essentially the same form as text. However, the language model has been optimized to operate on discrete, semantically meaningful tokens, while prevailing visual projectors transform an image into a long stream of continuous and highly correlated embeddings. This causes the visual tokens to behave differently from the word-like units that LLMs are originally trained to understand. We propose a novel Disentangled Visual Tokenization (DiVT) that clusters patch embeddings into coherent semantic units, so each token corresponds to a distinct visual concept instead of a rigid grid cell. DiVT further adapts its token budget to image complexity, providing an explicit accuracy-compute trade-off modifying neither the vision encoder nor the language model. Across diverse multimodal benchmarks, DiVT matches or surpasses baselines with significantly fewer visual tokens, demonstrating robustness under limited token budgets, significantly reducing memory cost and latency while making visual inputs more compatible with LLMs. Our code is available at https://github.com/snuviplab/DiVT.
title A More Word-like Image Tokenization for MLLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.17954