Unifying Vision-Language Latents for Zero-label Image Caption Enhancement

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Byun, Sanghyun, Guack, Jung Ick, Odema, Mohanad, Lee, Baisub, Song, Jacob, Chung, Woo Seong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908593643585536
author Byun, Sanghyun
Guack, Jung Ick
Odema, Mohanad
Lee, Baisub
Song, Jacob
Chung, Woo Seong
author_facet Byun, Sanghyun
Guack, Jung Ick
Odema, Mohanad
Lee, Baisub
Song, Jacob
Chung, Woo Seong
contents Vision-language models (VLMs) achieve remarkable performance through large-scale image-text pretraining. However, their reliance on labeled image datasets limits scalability and leaves vast amounts of unlabeled image data underutilized. To address this, we propose Unified Vision-Language Alignment for Zero-Label Enhancement (ViZer), an enhancement training framework that enables zero-label learning in image captioning, providing a practical starting point for broader zero-label adaptation in vision-language tasks. Unlike prior approaches that rely on human or synthetically annotated datasets, ViZer actively aligns vision and language representation features during training, enabling existing VLMs to generate improved captions without requiring text labels or full retraining. We demonstrate ViZer's advantage in qualitative evaluation, as automated caption metrics such as CIDEr and BERTScore often penalize details that are absent in reference captions. Applying ViZer on SmolVLM-Base and Qwen2-VL, we observe consistent qualitative improvements, producing captions that are more grounded and descriptive than their baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12931
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
Byun, Sanghyun
Guack, Jung Ick
Odema, Mohanad
Lee, Baisub
Song, Jacob
Chung, Woo Seong
Computer Vision and Pattern Recognition
Computation and Language
Vision-language models (VLMs) achieve remarkable performance through large-scale image-text pretraining. However, their reliance on labeled image datasets limits scalability and leaves vast amounts of unlabeled image data underutilized. To address this, we propose Unified Vision-Language Alignment for Zero-Label Enhancement (ViZer), an enhancement training framework that enables zero-label learning in image captioning, providing a practical starting point for broader zero-label adaptation in vision-language tasks. Unlike prior approaches that rely on human or synthetically annotated datasets, ViZer actively aligns vision and language representation features during training, enabling existing VLMs to generate improved captions without requiring text labels or full retraining. We demonstrate ViZer's advantage in qualitative evaluation, as automated caption metrics such as CIDEr and BERTScore often penalize details that are absent in reference captions. Applying ViZer on SmolVLM-Base and Qwen2-VL, we observe consistent qualitative improvements, producing captions that are more grounded and descriptive than their baseline.
title Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2510.12931