On the test-time zero-shot generalization of vision-language models: Do we really need prompt learning?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zanella, Maxime, Ayed, Ismail Ben |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Vocabulary-free few-shot learning for Vision-Language Models
par: Zanella, Maxime, et autres
Publié: (2025)
par: Zanella, Maxime, et autres
Publié: (2025)
Low-Rank Few-Shot Adaptation of Vision-Language Models
par: Zanella, Maxime, et autres
Publié: (2024)
par: Zanella, Maxime, et autres
Publié: (2024)
Boosting Vision-Language Models with Transduction
par: Zanella, Maxime, et autres
Publié: (2024)
par: Zanella, Maxime, et autres
Publié: (2024)
Language-Aware Information Maximization for Transductive Few-Shot CLIP
par: Baklouti, Ghassen, et autres
Publié: (2025)
par: Baklouti, Ghassen, et autres
Publié: (2025)
Realistic Test-Time Adaptation of Vision-Language Models
par: Zanella, Maxime, et autres
Publié: (2025)
par: Zanella, Maxime, et autres
Publié: (2025)
Boosting Vision-Language Models for Histopathology Classification: Predict all at once
par: Zanella, Maxime, et autres
Publié: (2024)
par: Zanella, Maxime, et autres
Publié: (2024)
bi-modal textual prompt learning for vision-language models in remote sensing
par: Kashyap, Pankhi, et autres
Publié: (2026)
par: Kashyap, Pankhi, et autres
Publié: (2026)
Are vision-language models ready to zero-shot replace supervised classification models in agriculture?
par: Ranario, Earl, et autres
Publié: (2025)
par: Ranario, Earl, et autres
Publié: (2025)
VCP-CLIP: A visual context prompting model for zero-shot anomaly segmentation
par: Qu, Zhen, et autres
Publié: (2024)
par: Qu, Zhen, et autres
Publié: (2024)
Zero-shot large vision-language model prompting for automated bone identification in paleoradiology x-ray archives
par: Dong, Owen, et autres
Publié: (2026)
par: Dong, Owen, et autres
Publié: (2026)
Are foundation models for computer vision good conformal predictors?
par: Fillioux, Leo, et autres
Publié: (2024)
par: Fillioux, Leo, et autres
Publié: (2024)
Few-shot Adaptation of Medical Vision-Language Models
par: Shakeri, Fereshteh, et autres
Publié: (2024)
par: Shakeri, Fereshteh, et autres
Publié: (2024)
Enhancing Remote Sensing Vision-Language Models for Zero-Shot Scene Classification
par: Khoury, Karim El, et autres
Publié: (2024)
par: Khoury, Karim El, et autres
Publié: (2024)
Pay Attention to Your Neighbours: Training-Free Open-Vocabulary Semantic Segmentation
par: Hajimiri, Sina, et autres
Publié: (2024)
par: Hajimiri, Sina, et autres
Publié: (2024)
ORION: ORthonormal Text Encoding for Universal VLM AdaptatION
par: Chakraborty, Omprakash, et autres
Publié: (2026)
par: Chakraborty, Omprakash, et autres
Publié: (2026)
InfraDiffusion: zero-shot depth map restoration with diffusion models and prompted segmentation from sparse infrastructure point clouds
par: Jing, Yixiong, et autres
Publié: (2025)
par: Jing, Yixiong, et autres
Publié: (2025)
SenCLIP: Enhancing zero-shot land-use mapping for Sentinel-2 with ground-level prompting
par: Jain, Pallavi, et autres
Publié: (2024)
par: Jain, Pallavi, et autres
Publié: (2024)
A Reality Check of Vision-Language Pre-training in Radiology: Have We Progressed Using Text?
par: Silva-Rodríguez, Julio, et autres
Publié: (2025)
par: Silva-Rodríguez, Julio, et autres
Publié: (2025)
Towards Foundation Models and Few-Shot Parameter-Efficient Fine-Tuning for Volumetric Organ Segmentation
par: Silva-Rodríguez, Julio, et autres
Publié: (2023)
par: Silva-Rodríguez, Julio, et autres
Publié: (2023)
Trustworthy Few-Shot Transfer of Medical VLMs through Split Conformal Prediction
par: Silva-Rodríguez, Julio, et autres
Publié: (2025)
par: Silva-Rodríguez, Julio, et autres
Publié: (2025)
Conformal Prediction for Zero-Shot Models
par: Silva-Rodríguez, Julio, et autres
Publié: (2025)
par: Silva-Rodríguez, Julio, et autres
Publié: (2025)
Do not trust what you trust: Miscalibration in Semi-supervised Learning
par: Mishra, Shambhavi, et autres
Publié: (2024)
par: Mishra, Shambhavi, et autres
Publié: (2024)
Initialization matters in few-shot adaptation of vision-language models for histopathological image classification
par: Meseguer, Pablo, et autres
Publié: (2026)
par: Meseguer, Pablo, et autres
Publié: (2026)
Do large language vision models understand 3D shapes?
par: Eppel, Sagi
Publié: (2024)
par: Eppel, Sagi
Publié: (2024)
Zero-shot segmentation of skin tumors in whole-slide images with vision-language foundation models
par: Moreno, Santiago, et autres
Publié: (2025)
par: Moreno, Santiago, et autres
Publié: (2025)
Noise-aware few-shot learning through bi-directional multi-view prompt alignment
par: Niu, Lu, et autres
Publié: (2026)
par: Niu, Lu, et autres
Publié: (2026)
Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection
par: Yang, Yuguang, et autres
Publié: (2025)
par: Yang, Yuguang, et autres
Publié: (2025)
Class and Region-Adaptive Constraints for Network Calibration
par: Murugesan, Balamurali, et autres
Publié: (2024)
par: Murugesan, Balamurali, et autres
Publié: (2024)
Robust Calibration of Large Vision-Language Adapters
par: Murugesan, Balamurali, et autres
Publié: (2024)
par: Murugesan, Balamurali, et autres
Publié: (2024)
A Closer Look at the Few-Shot Adaptation of Large Vision-Language Models
par: Silva-Rodríguez, Julio, et autres
Publié: (2023)
par: Silva-Rodríguez, Julio, et autres
Publié: (2023)
FreeZe: Training-free zero-shot 6D pose estimation with geometric and vision foundation models
par: Caraffa, Andrea, et autres
Publié: (2023)
par: Caraffa, Andrea, et autres
Publié: (2023)
Trust your neighbours: Penalty-based constraints for model calibration
par: Murugesan, Balamurali, et autres
Publié: (2023)
par: Murugesan, Balamurali, et autres
Publié: (2023)
Prompting classes: Exploring the Power of Prompt Class Learning in Weakly Supervised Semantic Segmentation
par: Murugesan, Balamurali, et autres
Publié: (2023)
par: Murugesan, Balamurali, et autres
Publié: (2023)
Calibrating Segmentation Networks with Margin-based Label Smoothing
par: Murugesan, Balamurali, et autres
Publié: (2022)
par: Murugesan, Balamurali, et autres
Publié: (2022)
A novel spatial-frequency domain network for zero-shot incremental learning
par: Ren, Jie, et autres
Publié: (2024)
par: Ren, Jie, et autres
Publié: (2024)
Sparsity Outperforms Low-Rank Projections in Few-Shot Adaptation
par: Mrabah, Nairouz, et autres
Publié: (2025)
par: Mrabah, Nairouz, et autres
Publié: (2025)
Semantic Anchor Transport: Robust Test-Time Adaptation for Vision-Language Models
par: Mishra, Shambhavi, et autres
Publié: (2024)
par: Mishra, Shambhavi, et autres
Publié: (2024)
Few-Shot, Now for Real: Medical VLMs Adaptation without Balanced Sets or Validation
par: Silva-Rodríguez, Julio, et autres
Publié: (2025)
par: Silva-Rodríguez, Julio, et autres
Publié: (2025)
A Foundation Language-Image Model of the Retina (FLAIR): Encoding Expert Knowledge in Text Supervision
par: Silva-Rodríguez, Julio, et autres
Publié: (2023)
par: Silva-Rodríguez, Julio, et autres
Publié: (2023)
Regularized Low-Rank Adaptation for Few-Shot Organ Segmentation
par: Baklouti, Ghassen, et autres
Publié: (2025)
par: Baklouti, Ghassen, et autres
Publié: (2025)
Documents similaires
-
Vocabulary-free few-shot learning for Vision-Language Models
par: Zanella, Maxime, et autres
Publié: (2025) -
Low-Rank Few-Shot Adaptation of Vision-Language Models
par: Zanella, Maxime, et autres
Publié: (2024) -
Boosting Vision-Language Models with Transduction
par: Zanella, Maxime, et autres
Publié: (2024) -
Language-Aware Information Maximization for Transductive Few-Shot CLIP
par: Baklouti, Ghassen, et autres
Publié: (2025) -
Realistic Test-Time Adaptation of Vision-Language Models
par: Zanella, Maxime, et autres
Publié: (2025)