Saved in:
Bibliographic Details
Main Authors: Wei, Xiaoyang, Kurtz, Camille, Cloppet, Florence
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.13876
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912715485741056
author Wei, Xiaoyang
Kurtz, Camille
Cloppet, Florence
author_facet Wei, Xiaoyang
Kurtz, Camille
Cloppet, Florence
contents Contrastive Language-Image Pretraining (CLIP) has demonstrated strong generalization for vision-language tasks in computer vision and medical domains, yet its text encoder accepts only up to 77 tokens, which limits its ability to represent long and information-rich radiology reports. Recent adaptations using domain-specific encoders, such as PubMedBERT or ClinicalBERT, mitigate this issue by leveraging medical corpora, but remain constrained by their limited input length (typically 512 tokens) and relatively shallow semantic understanding. To address these limitations, we propose QwenCLIP, a vision-language framework that replaces CLIP's text encoder with a large language model (LLM)-based embedding module (e.g., Qwen3-Embedding) and introduces learnable prompts to enhance cross-modal alignment. By leveraging the extended context window and richer representations of LLMs, QwenCLIP captures comprehensive medical semantics from long-form clinical text, substantially improving medical image-text alignment and downstream performance on radiology benchmarks. Our code is publicly available at https://github.com/Wxy-24/QwenCLIP.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13876
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QwenCLIP: Boosting Medical Vision-Language Pretraining via LLM Embeddings and Prompt tuning
Wei, Xiaoyang
Kurtz, Camille
Cloppet, Florence
Computer Vision and Pattern Recognition
Contrastive Language-Image Pretraining (CLIP) has demonstrated strong generalization for vision-language tasks in computer vision and medical domains, yet its text encoder accepts only up to 77 tokens, which limits its ability to represent long and information-rich radiology reports. Recent adaptations using domain-specific encoders, such as PubMedBERT or ClinicalBERT, mitigate this issue by leveraging medical corpora, but remain constrained by their limited input length (typically 512 tokens) and relatively shallow semantic understanding. To address these limitations, we propose QwenCLIP, a vision-language framework that replaces CLIP's text encoder with a large language model (LLM)-based embedding module (e.g., Qwen3-Embedding) and introduces learnable prompts to enhance cross-modal alignment. By leveraging the extended context window and richer representations of LLMs, QwenCLIP captures comprehensive medical semantics from long-form clinical text, substantially improving medical image-text alignment and downstream performance on radiology benchmarks. Our code is publicly available at https://github.com/Wxy-24/QwenCLIP.
title QwenCLIP: Boosting Medical Vision-Language Pretraining via LLM Embeddings and Prompt tuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.13876