Saved in:
| Main Authors: | Wei, Xiaoyang, Kurtz, Camille, Cloppet, Florence |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.13876 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Semantically-Aware Relevance Measure for Content-Based Medical Image Retrieval Evaluation
by: Wei, Xiaoyang, et al.
Published: (2025)
by: Wei, Xiaoyang, et al.
Published: (2025)
Prompt Tuning for CLIP on the Pretrained Manifold
by: Yang, Xi, et al.
Published: (2026)
by: Yang, Xi, et al.
Published: (2026)
Revisiting Prompt Pretraining of Vision-Language Models
by: Chen, Zhenyuan, et al.
Published: (2024)
by: Chen, Zhenyuan, et al.
Published: (2024)
Pedestrian Attribute Recognition via CLIP based Prompt Vision-Language Fusion
by: Wang, Xiao, et al.
Published: (2023)
by: Wang, Xiao, et al.
Published: (2023)
Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
by: Pham, Phuc, et al.
Published: (2025)
by: Pham, Phuc, et al.
Published: (2025)
Benchmarking Pretrained Vision Embeddings for Near- and Duplicate Detection in Medical Images
by: Truong, Tuan, et al.
Published: (2023)
by: Truong, Tuan, et al.
Published: (2023)
MedP-CLIP: Medical CLIP with Region-Aware Prompt Integration
by: Peng, Jiahui, et al.
Published: (2026)
by: Peng, Jiahui, et al.
Published: (2026)
Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
by: Liu, Yufang, et al.
Published: (2024)
by: Liu, Yufang, et al.
Published: (2024)
Language Prompt vs. Image Enhancement: Boosting Object Detection With CLIP in Hazy Environments
by: Pang, Jian, et al.
Published: (2026)
by: Pang, Jian, et al.
Published: (2026)
ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
by: Hu, Xiaoxing, et al.
Published: (2025)
by: Hu, Xiaoxing, et al.
Published: (2025)
MaskedCLIP: Bridging the Masked and CLIP Space for Semi-Supervised Medical Vision-Language Pre-training
by: Zhu, Lei, et al.
Published: (2025)
by: Zhu, Lei, et al.
Published: (2025)
NEARL-CLIP: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding
by: Peng, Zelin, et al.
Published: (2025)
by: Peng, Zelin, et al.
Published: (2025)
InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment Transfer
by: Yuan, Muyao, et al.
Published: (2025)
by: Yuan, Muyao, et al.
Published: (2025)
MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration
by: Zhang, Chenran, et al.
Published: (2026)
by: Zhang, Chenran, et al.
Published: (2026)
MAM-CLIP: Vision-Language Pretraining on Mammography Atlases for BI-RADS Classification
by: Gulluk, Halil Ibrahim, et al.
Published: (2026)
by: Gulluk, Halil Ibrahim, et al.
Published: (2026)
Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings
by: Ospanov, Azim, et al.
Published: (2024)
by: Ospanov, Azim, et al.
Published: (2024)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
by: Zhu, Wencheng, et al.
Published: (2025)
by: Zhu, Wencheng, et al.
Published: (2025)
Improving Calibration in Test-Time Prompt Tuning for Vision-Language Models via Data-Free Flatness-Aware Prompt Pretraining
by: Jang, Hyeonseo, et al.
Published: (2026)
by: Jang, Hyeonseo, et al.
Published: (2026)
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
by: Li, Yunheng, et al.
Published: (2024)
by: Li, Yunheng, et al.
Published: (2024)
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
FALIP: Visual Prompt as Foveal Attention Boosts CLIP Zero-Shot Performance
by: Zhuang, Jiedong, et al.
Published: (2024)
by: Zhuang, Jiedong, et al.
Published: (2024)
Conflict Adaptation in Vision-Language Models
by: Hu, Xiaoyang
Published: (2025)
by: Hu, Xiaoyang
Published: (2025)
LLM-Guided Diagnostic Evidence Alignment for Medical Vision-Language Pretraining under Limited Pairing
by: Yan, Huimin, et al.
Published: (2026)
by: Yan, Huimin, et al.
Published: (2026)
Envisioning MedCLIP: A Deep Dive into Explainability for Medical Vision-Language Models
by: Hashmi, Anees Ur Rehman, et al.
Published: (2024)
by: Hashmi, Anees Ur Rehman, et al.
Published: (2024)
PureCLIP-Depth: Prompt-Free and Decoder-Free Monocular Depth Estimation within CLIP Embedding Space
by: Miya, Ryutaro, et al.
Published: (2026)
by: Miya, Ryutaro, et al.
Published: (2026)
CLIP-Mamba: CLIP Pretrained Mamba Models with OOD and Hessian Evaluation
by: Huang, Weiquan, et al.
Published: (2024)
by: Huang, Weiquan, et al.
Published: (2024)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
QwenStyle: Content-Preserving Style Transfer with Qwen-Image-Edit
by: Zhang, Shiwen, et al.
Published: (2026)
by: Zhang, Shiwen, et al.
Published: (2026)
Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping
by: He, Dexuan, et al.
Published: (2025)
by: He, Dexuan, et al.
Published: (2025)
BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIP
by: Bai, Jiawang, et al.
Published: (2023)
by: Bai, Jiawang, et al.
Published: (2023)
Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding
by: Yao, Yuan, et al.
Published: (2026)
by: Yao, Yuan, et al.
Published: (2026)
Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
by: Jin, Yang, et al.
Published: (2023)
by: Jin, Yang, et al.
Published: (2023)
FrEVL: Leveraging Frozen Pretrained Embeddings for Efficient Vision-Language Understanding
by: Bourigault, Emmanuelle, et al.
Published: (2025)
by: Bourigault, Emmanuelle, et al.
Published: (2025)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining
by: Hu, Ming, et al.
Published: (2024)
by: Hu, Ming, et al.
Published: (2024)
AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models
by: Metzen, Jan Hendrik, et al.
Published: (2023)
by: Metzen, Jan Hendrik, et al.
Published: (2023)
PromptSmooth: Certifying Robustness of Medical Vision-Language Models via Prompt Learning
by: Hussein, Noor, et al.
Published: (2024)
by: Hussein, Noor, et al.
Published: (2024)
Calibration-Aware Prompt Learning for Medical Vision-Language Models
by: Basu, Abhishek, et al.
Published: (2025)
by: Basu, Abhishek, et al.
Published: (2025)
PhenoLIP: Integrating Phenotype Ontology Knowledge into Medical Vision-Language Pretraining
by: Liang, Cheng, et al.
Published: (2026)
by: Liang, Cheng, et al.
Published: (2026)
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
by: Huang, Dawei, et al.
Published: (2025)
by: Huang, Dawei, et al.
Published: (2025)
Similar Items
-
A Semantically-Aware Relevance Measure for Content-Based Medical Image Retrieval Evaluation
by: Wei, Xiaoyang, et al.
Published: (2025) -
Prompt Tuning for CLIP on the Pretrained Manifold
by: Yang, Xi, et al.
Published: (2026) -
Revisiting Prompt Pretraining of Vision-Language Models
by: Chen, Zhenyuan, et al.
Published: (2024) -
Pedestrian Attribute Recognition via CLIP based Prompt Vision-Language Fusion
by: Wang, Xiao, et al.
Published: (2023) -
Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
by: Pham, Phuc, et al.
Published: (2025)