Saved in:
Bibliographic Details
Main Authors: Giannone, Giorgio, Li, Ruoteng, Feng, Qianli, Perevodchikov, Evgeny, Chen, Rui, Martinez, Aleix
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2501.04568
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913846471426048
author Giannone, Giorgio
Li, Ruoteng
Feng, Qianli
Perevodchikov, Evgeny
Chen, Rui
Martinez, Aleix
author_facet Giannone, Giorgio
Li, Ruoteng
Feng, Qianli
Perevodchikov, Evgeny
Chen, Rui
Martinez, Aleix
contents Vision-language models (VLMs) have demonstrated remarkable potential in integrating visual and linguistic information, but their performance is often constrained by the need for extensive, high-quality image-text training data. Curation of these image-text pairs is both time-consuming and computationally expensive. To address this challenge, we introduce SVP (Sampling-based Visual Projection), a novel framework that enhances vision-language alignment without relying on manually curated text-image pairs or preference annotation. SVP leverages a small set of manually selected images, self-captioning and a pre-trained grounding model as a feedback mechanism to elicit latent information in VLMs. We evaluate our approach across six key areas: captioning, referring, visual question answering, multitasking, hallucination control, and object recall. Results demonstrate significant improvements, including a 14 % average improvement in captioning tasks, up to 12 % increase in object recall, and significantly reduced hallucinations, while maintaining question-answering capabilities. Using SVP, a small VLM achieves hallucination reductions similar to a model five times larger, while a VLM with initially poor referring capabilities more than doubles its performance, approaching parity with a model twice its size.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04568
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Feedback-Driven Vision-Language Alignment with Minimal Human Supervision
Giannone, Giorgio
Li, Ruoteng
Feng, Qianli
Perevodchikov, Evgeny
Chen, Rui
Martinez, Aleix
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Vision-language models (VLMs) have demonstrated remarkable potential in integrating visual and linguistic information, but their performance is often constrained by the need for extensive, high-quality image-text training data. Curation of these image-text pairs is both time-consuming and computationally expensive. To address this challenge, we introduce SVP (Sampling-based Visual Projection), a novel framework that enhances vision-language alignment without relying on manually curated text-image pairs or preference annotation. SVP leverages a small set of manually selected images, self-captioning and a pre-trained grounding model as a feedback mechanism to elicit latent information in VLMs. We evaluate our approach across six key areas: captioning, referring, visual question answering, multitasking, hallucination control, and object recall. Results demonstrate significant improvements, including a 14 % average improvement in captioning tasks, up to 12 % increase in object recall, and significantly reduced hallucinations, while maintaining question-answering capabilities. Using SVP, a small VLM achieves hallucination reductions similar to a model five times larger, while a VLM with initially poor referring capabilities more than doubles its performance, approaching parity with a model twice its size.
title Feedback-Driven Vision-Language Alignment with Minimal Human Supervision
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2501.04568