From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Mingxiao, Qu, Fang, Chen, Zhanpeng, Su, Na, Zhong, Zhizhou, Chen, Ziyang, Du, Nan, Li, Xiaolong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910824771092480
author Li, Mingxiao
Qu, Fang
Chen, Zhanpeng
Su, Na
Zhong, Zhizhou
Chen, Ziyang
Du, Nan
Li, Xiaolong
author_facet Li, Mingxiao
Qu, Fang
Chen, Zhanpeng
Su, Na
Zhong, Zhizhou
Chen, Ziyang
Du, Nan
Li, Xiaolong
contents While MLLMs perform well on perceptual tasks, they lack precise multimodal alignment, limiting performance. To address this challenge, we propose Vision Dynamic Embedding-Guided Pretraining (VDEP), a hybrid autoregressive training paradigm for MLLMs. Utilizing dynamic embeddings from the MLP following the visual encoder, this approach supervises image hidden states and integrates image tokens into autoregressive training. Existing MLLMs primarily focused on recovering information from textual inputs, often neglecting the effective processing of image data. In contrast, the key improvement of this work is the reinterpretation of multimodal alignment as a process of recovering information from input data, with particular emphasis on reconstructing detailed visual features.The proposed method seamlessly integrates into standard models without architectural changes. Experiments on 13 benchmarks show VDEP outperforms baselines, surpassing existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2502_09093
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs
Li, Mingxiao
Qu, Fang
Chen, Zhanpeng
Su, Na
Zhong, Zhizhou
Chen, Ziyang
Du, Nan
Li, Xiaolong
Computer Vision and Pattern Recognition
While MLLMs perform well on perceptual tasks, they lack precise multimodal alignment, limiting performance. To address this challenge, we propose Vision Dynamic Embedding-Guided Pretraining (VDEP), a hybrid autoregressive training paradigm for MLLMs. Utilizing dynamic embeddings from the MLP following the visual encoder, this approach supervises image hidden states and integrates image tokens into autoregressive training. Existing MLLMs primarily focused on recovering information from textual inputs, often neglecting the effective processing of image data. In contrast, the key improvement of this work is the reinterpretation of multimodal alignment as a process of recovering information from input data, with particular emphasis on reconstructing detailed visual features.The proposed method seamlessly integrates into standard models without architectural changes. Experiments on 13 benchmarks show VDEP outperforms baselines, surpassing existing methods.
title From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.09093