VladVA: Discriminative Fine-tuning of LVLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ouali, Yassine, Bulat, Adrian, Xenos, Alexandros, Zaganidis, Anestis, Metaxas, Ioannis Maniadis, Martinez, Brais, Tzimiropoulos, Georgios
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916727712907264
author Ouali, Yassine
Bulat, Adrian
Xenos, Alexandros
Zaganidis, Anestis
Metaxas, Ioannis Maniadis
Martinez, Brais
Tzimiropoulos, Georgios
author_facet Ouali, Yassine
Bulat, Adrian
Xenos, Alexandros
Zaganidis, Anestis
Metaxas, Ioannis Maniadis
Martinez, Brais
Tzimiropoulos, Georgios
contents Contrastively-trained Vision-Language Models (VLMs) like CLIP have become the de facto approach for discriminative vision-language representation learning. However, these models have limited language understanding, often exhibiting a "bag of words" behavior. At the same time, Large Vision-Language Models (LVLMs), which combine vision encoders with LLMs, have been shown to be capable of detailed vision-language reasoning, yet their autoregressive nature renders them less suitable for discriminative tasks. In this work, we propose to combine "the best of both worlds": a new training approach for discriminative fine-tuning of LVLMs that results in strong discriminative and compositional capabilities. Essentially, our approach converts a generative LVLM into a discriminative one, unlocking its capability for powerful image-text discrimination combined with enhanced language understanding. Our contributions include (1) a carefully designed training/optimization framework that utilizes image-text pairs of variable length and granularity for training the model with both contrastive and next-token prediction losses. This is accompanied by ablation studies that justify the necessity of our framework's components; (2) a parameter-efficient adaptation method using a combination of soft prompting and LoRA adapters; (3) significant improvements over state-of-the-art CLIP-like models of similar size, including standard image-text retrieval benchmarks and notable gains in compositionality.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04378
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VladVA: Discriminative Fine-tuning of LVLMs
Ouali, Yassine
Bulat, Adrian
Xenos, Alexandros
Zaganidis, Anestis
Metaxas, Ioannis Maniadis
Martinez, Brais
Tzimiropoulos, Georgios
Computer Vision and Pattern Recognition
Artificial Intelligence
Contrastively-trained Vision-Language Models (VLMs) like CLIP have become the de facto approach for discriminative vision-language representation learning. However, these models have limited language understanding, often exhibiting a "bag of words" behavior. At the same time, Large Vision-Language Models (LVLMs), which combine vision encoders with LLMs, have been shown to be capable of detailed vision-language reasoning, yet their autoregressive nature renders them less suitable for discriminative tasks. In this work, we propose to combine "the best of both worlds": a new training approach for discriminative fine-tuning of LVLMs that results in strong discriminative and compositional capabilities. Essentially, our approach converts a generative LVLM into a discriminative one, unlocking its capability for powerful image-text discrimination combined with enhanced language understanding. Our contributions include (1) a carefully designed training/optimization framework that utilizes image-text pairs of variable length and granularity for training the model with both contrastive and next-token prediction losses. This is accompanied by ablation studies that justify the necessity of our framework's components; (2) a parameter-efficient adaptation method using a combination of soft prompting and LoRA adapters; (3) significant improvements over state-of-the-art CLIP-like models of similar size, including standard image-text retrieval benchmarks and notable gains in compositionality.
title VladVA: Discriminative Fine-tuning of LVLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.04378