ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Yin, Yang, Kaicheng, Liang, Peirou, An, Xiang, Zhao, Yongle, Wang, Yumeng, Feng, Ziyong, Miles, Roy, Elezi, Ismail, Deng, Jiankang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916894513037312
author Xie, Yin
Yang, Kaicheng
Liang, Peirou
An, Xiang
Zhao, Yongle
Wang, Yumeng
Feng, Ziyong
Miles, Roy
Elezi, Ismail
Deng, Jiankang
author_facet Xie, Yin
Yang, Kaicheng
Liang, Peirou
An, Xiang
Zhao, Yongle
Wang, Yumeng
Feng, Ziyong
Miles, Roy
Elezi, Ismail
Deng, Jiankang
contents Large Multimodal Models (LMMs) often face a modality representation gap during pretraining: while language embeddings remain stable, visual representations are highly sensitive to contextual noise (e.g., background clutter). To address this issue, we introduce a visual comprehension stage, which we call ViCToR (Visual Comprehension via Token Reconstruction), a novel pretraining framework for LMMs. ViCToR employs a learnable visual token pool and utilizes the Hungarian matching algorithm to select semantically relevant tokens from this pool for visual token replacement. Furthermore, by integrating a visual token reconstruction loss with dense semantic supervision, ViCToR can learn tokens which retain high visual detail, thereby enhancing the large language model's (LLM's) understanding of visual information. After pretraining on 3 million publicly accessible images and captions, ViCToR achieves state-of-the-art results, improving over LLaVA-NeXT-8B by 10.4%, 3.2%, and 7.2% on the MMStar, SEED$^I$, and RealWorldQA benchmarks, respectively. Code is available at https://github.com/deepglint/Victor.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14332
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
Xie, Yin
Yang, Kaicheng
Liang, Peirou
An, Xiang
Zhao, Yongle
Wang, Yumeng
Feng, Ziyong
Miles, Roy
Elezi, Ismail
Deng, Jiankang
Computer Vision and Pattern Recognition
Large Multimodal Models (LMMs) often face a modality representation gap during pretraining: while language embeddings remain stable, visual representations are highly sensitive to contextual noise (e.g., background clutter). To address this issue, we introduce a visual comprehension stage, which we call ViCToR (Visual Comprehension via Token Reconstruction), a novel pretraining framework for LMMs. ViCToR employs a learnable visual token pool and utilizes the Hungarian matching algorithm to select semantically relevant tokens from this pool for visual token replacement. Furthermore, by integrating a visual token reconstruction loss with dense semantic supervision, ViCToR can learn tokens which retain high visual detail, thereby enhancing the large language model's (LLM's) understanding of visual information. After pretraining on 3 million publicly accessible images and captions, ViCToR achieves state-of-the-art results, improving over LLaVA-NeXT-8B by 10.4%, 3.2%, and 7.2% on the MMStar, SEED$^I$, and RealWorldQA benchmarks, respectively. Code is available at https://github.com/deepglint/Victor.
title ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.14332