Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Qiao, Liu, Yanjiang, Zhou, Weixiang, He, Ben, Lu, Yaojie, Lin, Hongyu, Zheng, Jia, Han, Xianpei, Sun, Le, Sun, Yingfei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909628011380736
author Liang, Qiao
Liu, Yanjiang
Zhou, Weixiang
He, Ben
Lu, Yaojie
Lin, Hongyu
Zheng, Jia
Han, Xianpei
Sun, Le
Sun, Yingfei
author_facet Liang, Qiao
Liu, Yanjiang
Zhou, Weixiang
He, Ben
Lu, Yaojie
Lin, Hongyu
Zheng, Jia
Han, Xianpei
Sun, Le
Sun, Yingfei
contents Does the prior knowledge of the vision encoder constrain the capability boundary of Multi-modal Large Language Models (MLLMs)? While most existing research treats MLLMs as unified systems optimized through end-to-end training, the impact of vision encoder's prior knowledge is seldom investigated. In this work, we introduce a novel metric, $Rank_e$, to quantify the effect of prior knowledge of the vision encoder on MLLM performance. Our analysis reveals a positive correlation between prior knowledge and MLLM performance. Moreover, we find that domain-specific fine-tuning using solely end-to-end visual question answering (VQA) data is insufficient, particularly for entities with low inherent visual prior knowledge. To address this issue, we propose VisPRE (Vision Prior Remediation), a two-stage training framework that explicitly incorporates prior knowledge at the vision encoder level. Experimental results demonstrate that augmenting vision encoder's prior knowledge substantially boosts the visual understanding capabilities of MLLMs, offering a novel and effective strategy for improving performance, especially in scenarios involving uncommon visual entities.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18034
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
Liang, Qiao
Liu, Yanjiang
Zhou, Weixiang
He, Ben
Lu, Yaojie
Lin, Hongyu
Zheng, Jia
Han, Xianpei
Sun, Le
Sun, Yingfei
Computer Vision and Pattern Recognition
Computation and Language
Does the prior knowledge of the vision encoder constrain the capability boundary of Multi-modal Large Language Models (MLLMs)? While most existing research treats MLLMs as unified systems optimized through end-to-end training, the impact of vision encoder's prior knowledge is seldom investigated. In this work, we introduce a novel metric, $Rank_e$, to quantify the effect of prior knowledge of the vision encoder on MLLM performance. Our analysis reveals a positive correlation between prior knowledge and MLLM performance. Moreover, we find that domain-specific fine-tuning using solely end-to-end visual question answering (VQA) data is insufficient, particularly for entities with low inherent visual prior knowledge. To address this issue, we propose VisPRE (Vision Prior Remediation), a two-stage training framework that explicitly incorporates prior knowledge at the vision encoder level. Experimental results demonstrate that augmenting vision encoder's prior knowledge substantially boosts the visual understanding capabilities of MLLMs, offering a novel and effective strategy for improving performance, especially in scenarios involving uncommon visual entities.
title Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2503.18034