Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Mingyuan, Bai, Yue, Wang, Yifan, Huang, Yiyang, Fu, Yun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908736847609856
author Zhang, Mingyuan
Bai, Yue
Wang, Yifan
Huang, Yiyang
Fu, Yun
author_facet Zhang, Mingyuan
Bai, Yue
Wang, Yifan
Huang, Yiyang
Fu, Yun
contents Explorations in fine-tuning Vision-Language Models (VLMs), such as Low-Rank Adaptation (LoRA) from Parameter Efficient Fine-Tuning (PEFT), have made impressive progress. However, most approaches rely on explicit weight updates, overlooking the extensive representational structures already encoded in pre-trained models that remain underutilized. Recent works have demonstrated that Mask Fine-Tuning (MFT) can be a powerful and efficient post-training paradigm for language models. Instead of updating weights, MFT assigns learnable gating scores to each weight, allowing the model to reorganize its internal subnetworks for downstream task adaptation. In this paper, we rethink fine-tuning for VLMs from a structural reparameterization perspective grounded in MFT. We apply MFT to the language and projector components of VLMs with different language backbones and compare against strong PEFT baselines. Experiments show that MFT consistently surpasses LoRA variants and even full fine-tuning, achieving high performance without altering the frozen backbone. Our findings reveal that effective adaptation can emerge not only from updating weights but also from reestablishing connections among the model's existing knowledge. Code available at: https://github.com/Ming-K9/MFT-VLM
format Preprint
id arxiv_https___arxiv_org_abs_2512_23073
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
Zhang, Mingyuan
Bai, Yue
Wang, Yifan
Huang, Yiyang
Fu, Yun
Machine Learning
Computer Vision and Pattern Recognition
Explorations in fine-tuning Vision-Language Models (VLMs), such as Low-Rank Adaptation (LoRA) from Parameter Efficient Fine-Tuning (PEFT), have made impressive progress. However, most approaches rely on explicit weight updates, overlooking the extensive representational structures already encoded in pre-trained models that remain underutilized. Recent works have demonstrated that Mask Fine-Tuning (MFT) can be a powerful and efficient post-training paradigm for language models. Instead of updating weights, MFT assigns learnable gating scores to each weight, allowing the model to reorganize its internal subnetworks for downstream task adaptation. In this paper, we rethink fine-tuning for VLMs from a structural reparameterization perspective grounded in MFT. We apply MFT to the language and projector components of VLMs with different language backbones and compare against strong PEFT baselines. Experiments show that MFT consistently surpasses LoRA variants and even full fine-tuning, achieving high performance without altering the frozen backbone. Our findings reveal that effective adaptation can emerge not only from updating weights but also from reestablishing connections among the model's existing knowledge. Code available at: https://github.com/Ming-K9/MFT-VLM
title Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.23073