LongProLIP: A Probabilistic Vision-Language Model with Long Context Text

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chun, Sanghyuk, Yun, Sangdoo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929756997419008
author Chun, Sanghyuk
Yun, Sangdoo
author_facet Chun, Sanghyuk
Yun, Sangdoo
contents Recently, Probabilistic Language-Image Pre-Training (ProLIP) has been proposed to tackle the multiplicity issue of vision-language (VL) tasks. Despite their success in probabilistic representation learning at a scale, the ProLIP models cannot handle long context texts longer than 64 context length, which limits their ability to capture rich contextual information from longer text sequences. To address this issue, this paper proposes a fine-tuning strategy for ProLIP to accept longer texts, e.g., 256 text tokens. Experimental results on Urban-1k and the DataComp evaluation suite show that the proposed LongProLIP recipe can improve understanding of long contexts while minimizing the negative effect of fine-tuning.We also observe a trade-off between the long context understanding (measured by Urban-1k) and general zero-shot capability (measured by evaluation datasets by DataComp). Code is available at https://github.com/naver-ai/prolip
format Preprint
id arxiv_https___arxiv_org_abs_2503_08048
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
Chun, Sanghyuk
Yun, Sangdoo
Computer Vision and Pattern Recognition
Machine Learning
Recently, Probabilistic Language-Image Pre-Training (ProLIP) has been proposed to tackle the multiplicity issue of vision-language (VL) tasks. Despite their success in probabilistic representation learning at a scale, the ProLIP models cannot handle long context texts longer than 64 context length, which limits their ability to capture rich contextual information from longer text sequences. To address this issue, this paper proposes a fine-tuning strategy for ProLIP to accept longer texts, e.g., 256 text tokens. Experimental results on Urban-1k and the DataComp evaluation suite show that the proposed LongProLIP recipe can improve understanding of long contexts while minimizing the negative effect of fine-tuning.We also observe a trade-off between the long context understanding (measured by Urban-1k) and general zero-shot capability (measured by evaluation datasets by DataComp). Code is available at https://github.com/naver-ai/prolip
title LongProLIP: A Probabilistic Vision-Language Model with Long Context Text
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2503.08048