VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Yikun, Liu, Yuan, Di, Shangzhe, Wang, Haicheng, Zhao, Zhongyin, Tian, Le, Zhou, Xiao, Zhou, Jie, Yao, Jiangchao, Wang, Yanfeng, Xie, Weidi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911438174420992
author Liu, Yikun
Liu, Yuan
Di, Shangzhe
Wang, Haicheng
Zhao, Zhongyin
Tian, Le
Zhou, Xiao
Zhou, Jie
Yao, Jiangchao
Wang, Yanfeng
Xie, Weidi
author_facet Liu, Yikun
Liu, Yuan
Di, Shangzhe
Wang, Haicheng
Zhao, Zhongyin
Tian, Le
Zhou, Xiao
Zhou, Jie
Yao, Jiangchao
Wang, Yanfeng
Xie, Weidi
contents Multimodal Large Language Models (MLLMs) have recently achieved remarkable success in visual-language understanding, demonstrating superior high-level semantic alignment within their vision encoders. An important question thus arises: Can these encoders serve as versatile vision backbones, capable of reliably performing classic vision-centric tasks as well? To address the question, we make the following contributions: (i) we identify that the vision encoders within MLLMs exhibit deficiencies in their dense feature representations, as evidenced by their suboptimal performance on dense prediction tasks (e.g., semantic segmentation, depth estimation); (ii) we propose VersaViT, a well-rounded vision transformer that instantiates a novel multi-task framework for collaborative post-training. This framework facilitates the optimization of the vision backbone via lightweight task heads with multi-granularity supervision; (iii) extensive experiments across various downstream tasks demonstrate the effectiveness of our method, yielding a versatile vision backbone suited for both language-mediated reasoning and pixel-level understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09934
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
Liu, Yikun
Liu, Yuan
Di, Shangzhe
Wang, Haicheng
Zhao, Zhongyin
Tian, Le
Zhou, Xiao
Zhou, Jie
Yao, Jiangchao
Wang, Yanfeng
Xie, Weidi
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) have recently achieved remarkable success in visual-language understanding, demonstrating superior high-level semantic alignment within their vision encoders. An important question thus arises: Can these encoders serve as versatile vision backbones, capable of reliably performing classic vision-centric tasks as well? To address the question, we make the following contributions: (i) we identify that the vision encoders within MLLMs exhibit deficiencies in their dense feature representations, as evidenced by their suboptimal performance on dense prediction tasks (e.g., semantic segmentation, depth estimation); (ii) we propose VersaViT, a well-rounded vision transformer that instantiates a novel multi-task framework for collaborative post-training. This framework facilitates the optimization of the vision backbone via lightweight task heads with multi-granularity supervision; (iii) extensive experiments across various downstream tasks demonstrate the effectiveness of our method, yielding a versatile vision backbone suited for both language-mediated reasoning and pixel-level understanding.
title VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.09934