Adapting Vision-Language Models Without Labels: A Comprehensive Survey
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Hao, Sheng, Lijun, Liang, Jian, He, Ran, Chatzi, Eleni, Fink, Olga |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
To Trust Or Not To Trust Your Vision-Language Model's Prediction
by: Dong, Hao, et al.
Published: (2025)
by: Dong, Hao, et al.
Published: (2025)
Multimodal Learning for Arcing Detection in Pantograph-Catenary Systems
by: Dong, Hao, et al.
Published: (2026)
by: Dong, Hao, et al.
Published: (2026)
Towards Multimodal Open-Set Domain Generalization and Adaptation through Self-supervision
by: Dong, Hao, et al.
Published: (2024)
by: Dong, Hao, et al.
Published: (2024)
Towards Robust Multimodal Open-set Test-time Adaptation via Adaptive Entropy-aware Optimization
by: Dong, Hao, et al.
Published: (2025)
by: Dong, Hao, et al.
Published: (2025)
MultiOOD: Scaling Out-of-Distribution Detection for Multiple Modalities
by: Dong, Hao, et al.
Published: (2024)
by: Dong, Hao, et al.
Published: (2024)
NNG-Mix: Improving Semi-supervised Anomaly Detection with Pseudo-anomaly Generation
by: Dong, Hao, et al.
Published: (2023)
by: Dong, Hao, et al.
Published: (2023)
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
by: Dong, Hao, et al.
Published: (2026)
by: Dong, Hao, et al.
Published: (2026)
Advances in Multimodal Adaptation and Generalization: From Traditional Approaches to Foundation Models
by: Dong, Hao, et al.
Published: (2025)
by: Dong, Hao, et al.
Published: (2025)
A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts
by: Liang, Jian, et al.
Published: (2023)
by: Liang, Jian, et al.
Published: (2023)
Recall and Refine: A Simple but Effective Source-free Open-set Domain Adaptation Framework
by: Nejjar, Ismail, et al.
Published: (2024)
by: Nejjar, Ismail, et al.
Published: (2024)
Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language Models
by: Wang, Zhengbo, et al.
Published: (2024)
by: Wang, Zhengbo, et al.
Published: (2024)
Semi-Supervised Semantic Segmentation Based on Pseudo-Labels: A Survey
by: Ran, Lingyan, et al.
Published: (2024)
by: Ran, Lingyan, et al.
Published: (2024)
The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection
by: Ai, Wei, et al.
Published: (2026)
by: Ai, Wei, et al.
Published: (2026)
Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving
by: Jiang, Hao, et al.
Published: (2025)
by: Jiang, Hao, et al.
Published: (2025)
A Hard-to-Beat Baseline for Training-free CLIP-based Adaptation
by: Wang, Zhengbo, et al.
Published: (2024)
by: Wang, Zhengbo, et al.
Published: (2024)
A Survey on Vision-Language-Action Models for Autonomous Driving
by: Jiang, Sicong, et al.
Published: (2025)
by: Jiang, Sicong, et al.
Published: (2025)
Adapting Vision-Language Models for E-commerce Understanding at Scale
by: Nulli, Matteo, et al.
Published: (2026)
by: Nulli, Matteo, et al.
Published: (2026)
LaViC: Adapting Large Vision-Language Models to Visually-Aware Conversational Recommendation
by: Jeon, Hyunsik, et al.
Published: (2025)
by: Jeon, Hyunsik, et al.
Published: (2025)
Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
by: Lian, Chenyu, et al.
Published: (2025)
by: Lian, Chenyu, et al.
Published: (2025)
Adapting Lightweight Vision Language Models for Radiological Visual Question Answering
by: Shourya, Aditya, et al.
Published: (2025)
by: Shourya, Aditya, et al.
Published: (2025)
Vision-Language Models for Edge Networks: A Comprehensive Survey
by: Sharshar, Ahmed, et al.
Published: (2025)
by: Sharshar, Ahmed, et al.
Published: (2025)
Mirror Target YOLO: An Improved YOLOv8 Method with Indirect Vision for Heritage Buildings Fire Detection
by: Liang, Jian, et al.
Published: (2024)
by: Liang, Jian, et al.
Published: (2024)
Adapting Vision-Language Models for Evaluating World Models
by: Hendriksen, Mariya, et al.
Published: (2025)
by: Hendriksen, Mariya, et al.
Published: (2025)
Aligned Vector Quantization for Edge-Cloud Collabrative Vision-Language Models
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
Adaptive Confidence Regularization for Multimodal Failure Detection
by: Liu, Moru, et al.
Published: (2026)
by: Liu, Moru, et al.
Published: (2026)
Sanitizing Manufacturing Dataset Labels Using Vision-Language Models
by: Mahjourian, Nazanin, et al.
Published: (2025)
by: Mahjourian, Nazanin, et al.
Published: (2025)
Tuning Vision-Language Models with Candidate Labels by Prompt Alignment
by: Zhang, Zhifang, et al.
Published: (2024)
by: Zhang, Zhifang, et al.
Published: (2024)
Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
by: Li, Jinxuan, et al.
Published: (2025)
by: Li, Jinxuan, et al.
Published: (2025)
Vision Transformers in Precision Agriculture: A Comprehensive Survey
by: Mehdipour, Saber, et al.
Published: (2025)
by: Mehdipour, Saber, et al.
Published: (2025)
MetaDent: Labeling Clinical Images for Vision-Language Models in Dentistry
by: Li, Meng-Xun, et al.
Published: (2026)
by: Li, Meng-Xun, et al.
Published: (2026)
R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning
by: Sheng, Lijun, et al.
Published: (2025)
by: Sheng, Lijun, et al.
Published: (2025)
The Illusion of Progress? A Critical Look at Test-Time Adaptation for Vision-Language Models
by: Sheng, Lijun, et al.
Published: (2025)
by: Sheng, Lijun, et al.
Published: (2025)
Improvise, Adapt, Overcome -- Telescopic Adapters for Efficient Fine-tuning of Vision Language Models in Medical Imaging
by: Mishra, Ujjwal, et al.
Published: (2025)
by: Mishra, Ujjwal, et al.
Published: (2025)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
by: Fu, Chaoyou, et al.
Published: (2024)
by: Fu, Chaoyou, et al.
Published: (2024)
DentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice
by: Meng, Zijie, et al.
Published: (2025)
by: Meng, Zijie, et al.
Published: (2025)
Automated Processing of eXplainable Artificial Intelligence Outputs in Deep Learning Models for Fault Diagnostics of Large Infrastructures
by: Floreale, Giovanni, et al.
Published: (2025)
by: Floreale, Giovanni, et al.
Published: (2025)
Sample Correlation for Fingerprinting Deep Face Recognition
by: Guan, Jiyang, et al.
Published: (2024)
by: Guan, Jiyang, et al.
Published: (2024)
Evaluating Vision Language Models (VLMs) for Radiology: A Comprehensive Analysis
by: Li, Frank, et al.
Published: (2025)
by: Li, Frank, et al.
Published: (2025)
Vision Language Models in Autonomous Driving: A Survey and Outlook
by: Zhou, Xingcheng, et al.
Published: (2023)
by: Zhou, Xingcheng, et al.
Published: (2023)
Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving
by: Yang, Sheng, et al.
Published: (2025)
by: Yang, Sheng, et al.
Published: (2025)
Similar Items
-
To Trust Or Not To Trust Your Vision-Language Model's Prediction
by: Dong, Hao, et al.
Published: (2025) -
Multimodal Learning for Arcing Detection in Pantograph-Catenary Systems
by: Dong, Hao, et al.
Published: (2026) -
Towards Multimodal Open-Set Domain Generalization and Adaptation through Self-supervision
by: Dong, Hao, et al.
Published: (2024) -
Towards Robust Multimodal Open-set Test-time Adaptation via Adaptive Entropy-aware Optimization
by: Dong, Hao, et al.
Published: (2025) -
MultiOOD: Scaling Out-of-Distribution Detection for Multiple Modalities
by: Dong, Hao, et al.
Published: (2024)