ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916894513037312 |
|---|---|
| author | Xie, Yin Yang, Kaicheng Liang, Peirou An, Xiang Zhao, Yongle Wang, Yumeng Feng, Ziyong Miles, Roy Elezi, Ismail Deng, Jiankang |
| author_facet | Xie, Yin Yang, Kaicheng Liang, Peirou An, Xiang Zhao, Yongle Wang, Yumeng Feng, Ziyong Miles, Roy Elezi, Ismail Deng, Jiankang |
| contents | Large Multimodal Models (LMMs) often face a modality representation gap during pretraining: while language embeddings remain stable, visual representations are highly sensitive to contextual noise (e.g., background clutter). To address this issue, we introduce a visual comprehension stage, which we call ViCToR (Visual Comprehension via Token Reconstruction), a novel pretraining framework for LMMs. ViCToR employs a learnable visual token pool and utilizes the Hungarian matching algorithm to select semantically relevant tokens from this pool for visual token replacement. Furthermore, by integrating a visual token reconstruction loss with dense semantic supervision, ViCToR can learn tokens which retain high visual detail, thereby enhancing the large language model's (LLM's) understanding of visual information. After pretraining on 3 million publicly accessible images and captions, ViCToR achieves state-of-the-art results, improving over LLaVA-NeXT-8B by 10.4%, 3.2%, and 7.2% on the MMStar, SEED$^I$, and RealWorldQA benchmarks, respectively. Code is available at https://github.com/deepglint/Victor. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_14332 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs Xie, Yin Yang, Kaicheng Liang, Peirou An, Xiang Zhao, Yongle Wang, Yumeng Feng, Ziyong Miles, Roy Elezi, Ismail Deng, Jiankang Computer Vision and Pattern Recognition Large Multimodal Models (LMMs) often face a modality representation gap during pretraining: while language embeddings remain stable, visual representations are highly sensitive to contextual noise (e.g., background clutter). To address this issue, we introduce a visual comprehension stage, which we call ViCToR (Visual Comprehension via Token Reconstruction), a novel pretraining framework for LMMs. ViCToR employs a learnable visual token pool and utilizes the Hungarian matching algorithm to select semantically relevant tokens from this pool for visual token replacement. Furthermore, by integrating a visual token reconstruction loss with dense semantic supervision, ViCToR can learn tokens which retain high visual detail, thereby enhancing the large language model's (LLM's) understanding of visual information. After pretraining on 3 million publicly accessible images and captions, ViCToR achieves state-of-the-art results, improving over LLaVA-NeXT-8B by 10.4%, 3.2%, and 7.2% on the MMStar, SEED$^I$, and RealWorldQA benchmarks, respectively. Code is available at https://github.com/deepglint/Victor. |
| title | ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2410.14332 |