Noise is an Efficient Learner for Zero-Shot Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Imam, Raza, Hanif, Asif, Zhang, Jian, Dawoud, Khaled Waleed, Kementchedjhieva, Yova, Yaqub, Mohammad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Robustness of Medical Vision-Language Models: Are they Truly Generalizable?
von: Imam, Raza, et al.
Veröffentlicht: (2025)
von: Imam, Raza, et al.
Veröffentlicht: (2025)
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
von: Huzaifa, Muhammad, et al.
Veröffentlicht: (2024)
von: Huzaifa, Muhammad, et al.
Veröffentlicht: (2024)
T3: Test-Time Model Merging in VLMs for Zero-Shot Medical Imaging Analysis
von: Imam, Raza, et al.
Veröffentlicht: (2025)
von: Imam, Raza, et al.
Veröffentlicht: (2025)
A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA
von: Shoer, Belal, et al.
Veröffentlicht: (2025)
von: Shoer, Belal, et al.
Veröffentlicht: (2025)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
von: Mohamed, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Mohamed, Abdelrahman, et al.
Veröffentlicht: (2025)
Decoupling Clinical and Class-Agnostic Features for Reliable Few-Shot Adaptation under Shift
von: Rahman, Umaima, et al.
Veröffentlicht: (2025)
von: Rahman, Umaima, et al.
Veröffentlicht: (2025)
Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models
von: Monon, Mashrafi, et al.
Veröffentlicht: (2026)
von: Monon, Mashrafi, et al.
Veröffentlicht: (2026)
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
von: Salazar, Israfel, et al.
Veröffentlicht: (2025)
von: Salazar, Israfel, et al.
Veröffentlicht: (2025)
Test-Time Low Rank Adaptation via Confidence Maximization for Zero-Shot Generalization of Vision-Language Models
von: Imam, Raza, et al.
Veröffentlicht: (2024)
von: Imam, Raza, et al.
Veröffentlicht: (2024)
Do Vision and Language Models Share Concepts? A Vector Space Alignment Study
von: Li, Jiaang, et al.
Veröffentlicht: (2023)
von: Li, Jiaang, et al.
Veröffentlicht: (2023)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2026)
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2026)
DARK: Diagonal-Anchored Repulsive Knowledge Distillation for Vision-Language Models under Extreme Compression
von: Saeed, Numan, et al.
Veröffentlicht: (2026)
von: Saeed, Numan, et al.
Veröffentlicht: (2026)
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
von: Shahgir, Haz Sameen, et al.
Veröffentlicht: (2026)
von: Shahgir, Haz Sameen, et al.
Veröffentlicht: (2026)
Can language-guided unsupervised adaptation improve medical image classification using unpaired images and texts?
von: Rahman, Umaima, et al.
Veröffentlicht: (2024)
von: Rahman, Umaima, et al.
Veröffentlicht: (2024)
FETAL-GAUGE: A Benchmark for Assessing Vision-Language Models in Fetal Ultrasound
von: Alasmawi, Hussain, et al.
Veröffentlicht: (2025)
von: Alasmawi, Hussain, et al.
Veröffentlicht: (2025)
LLMs Can Compensate for Deficiencies in Visual Representations
von: Takishita, Sho, et al.
Veröffentlicht: (2025)
von: Takishita, Sho, et al.
Veröffentlicht: (2025)
Open-Source Image Editing Models Are Zero-Shot Vision Learners
von: Liu, Wei, et al.
Veröffentlicht: (2026)
von: Liu, Wei, et al.
Veröffentlicht: (2026)
CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging
von: Imam, Raza, et al.
Veröffentlicht: (2024)
von: Imam, Raza, et al.
Veröffentlicht: (2024)
TransResNet: Integrating the Strengths of ViTs and CNNs for High Resolution Medical Image Segmentation via Feature Grafting
von: Sharif, Muhammad Hamza, et al.
Veröffentlicht: (2024)
von: Sharif, Muhammad Hamza, et al.
Veröffentlicht: (2024)
Envisioning MedCLIP: A Deep Dive into Explainability for Medical Vision-Language Models
von: Hashmi, Anees Ur Rehman, et al.
Veröffentlicht: (2024)
von: Hashmi, Anees Ur Rehman, et al.
Veröffentlicht: (2024)
Investigating Zero-Shot Diagnostic Pathology in Vision-Language Models with Efficient Prompt Design
von: Sharma, Vasudev, et al.
Veröffentlicht: (2025)
von: Sharma, Vasudev, et al.
Veröffentlicht: (2025)
Zero-shot World Models Are Developmentally Efficient Learners
von: Aw, Khai Loong, et al.
Veröffentlicht: (2026)
von: Aw, Khai Loong, et al.
Veröffentlicht: (2026)
Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
von: Lai, Yuxiang, et al.
Veröffentlicht: (2025)
von: Lai, Yuxiang, et al.
Veröffentlicht: (2025)
TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model
von: Li, Ao, et al.
Veröffentlicht: (2025)
von: Li, Ao, et al.
Veröffentlicht: (2025)
Pre-trained Vision and Language Transformers Are Few-Shot Incremental Learners
von: Park, Keon-Hee, et al.
Veröffentlicht: (2024)
von: Park, Keon-Hee, et al.
Veröffentlicht: (2024)
Making Large Vision Language Models to be Good Few-shot Learners
von: Liu, Fan, et al.
Veröffentlicht: (2024)
von: Liu, Fan, et al.
Veröffentlicht: (2024)
On Evaluating Adversarial Robustness of Volumetric Medical Segmentation Models
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024)
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024)
Training-Free Zero-Shot Temporal Action Detection with Vision-Language Models
von: Han, Chaolei, et al.
Veröffentlicht: (2025)
von: Han, Chaolei, et al.
Veröffentlicht: (2025)
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
von: Atabuzzaman, Md., et al.
Veröffentlicht: (2025)
von: Atabuzzaman, Md., et al.
Veröffentlicht: (2025)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
von: Chen, Shiming, et al.
Veröffentlicht: (2025)
von: Chen, Shiming, et al.
Veröffentlicht: (2025)
Learning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language Reasoning
von: Li, Rongjie, et al.
Veröffentlicht: (2024)
von: Li, Rongjie, et al.
Veröffentlicht: (2024)
Efficient and Context-Aware Label Propagation for Zero-/Few-Shot Training-Free Adaptation of Vision-Language Model
von: Li, Yushu, et al.
Veröffentlicht: (2024)
von: Li, Yushu, et al.
Veröffentlicht: (2024)
MMRINet: Efficient Mamba-Based Segmentation with Dual-Path Refinement for Low-Resource MRI Analysis
von: Elsayed, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Elsayed, Abdelrahman, et al.
Veröffentlicht: (2025)
Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation
von: Unmesh, Asim, et al.
Veröffentlicht: (2026)
von: Unmesh, Asim, et al.
Veröffentlicht: (2026)
Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
von: Yin, Xiaojie, et al.
Veröffentlicht: (2025)
von: Yin, Xiaojie, et al.
Veröffentlicht: (2025)
Zero-Shot Robustness of Vision Language Models Via Confidence-Aware Weighting
von: Naghavian, Nikoo, et al.
Veröffentlicht: (2025)
von: Naghavian, Nikoo, et al.
Veröffentlicht: (2025)
Exploring the Zero-Shot Capabilities of Vision-Language Models for Improving Gaze Following
von: Gupta, Anshul, et al.
Veröffentlicht: (2024)
von: Gupta, Anshul, et al.
Veröffentlicht: (2024)
Enhancing Remote Sensing Vision-Language Models for Zero-Shot Scene Classification
von: Khoury, Karim El, et al.
Veröffentlicht: (2024)
von: Khoury, Karim El, et al.
Veröffentlicht: (2024)
BaFTA: Backprop-Free Test-Time Adaptation For Zero-Shot Vision-Language Models
von: Hu, Xuefeng, et al.
Veröffentlicht: (2024)
von: Hu, Xuefeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On the Robustness of Medical Vision-Language Models: Are they Truly Generalizable?
von: Imam, Raza, et al.
Veröffentlicht: (2025) -
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
von: Huzaifa, Muhammad, et al.
Veröffentlicht: (2024) -
T3: Test-Time Model Merging in VLMs for Zero-Shot Medical Imaging Analysis
von: Imam, Raza, et al.
Veröffentlicht: (2025) -
A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA
von: Shoer, Belal, et al.
Veröffentlicht: (2025) -
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
von: Mohamed, Abdelrahman, et al.
Veröffentlicht: (2025)