A Hard-to-Beat Baseline for Training-free CLIP-based Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhengbo, Liang, Jian, Sheng, Lijun, He, Ran, Wang, Zilei, Tan, Tieniu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language Models
by: Wang, Zhengbo, et al.
Published: (2024)
by: Wang, Zhengbo, et al.
Published: (2024)
Realistic Unsupervised CLIP Fine-tuning with Universal Entropy Optimization
by: Liang, Jian, et al.
Published: (2023)
by: Liang, Jian, et al.
Published: (2023)
Towards Compatible Fine-tuning for Vision-Language Model Updates
by: Wang, Zhengbo, et al.
Published: (2024)
by: Wang, Zhengbo, et al.
Published: (2024)
The Illusion of Progress? A Critical Look at Test-Time Adaptation for Vision-Language Models
by: Sheng, Lijun, et al.
Published: (2025)
by: Sheng, Lijun, et al.
Published: (2025)
A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts
by: Liang, Jian, et al.
Published: (2023)
by: Liang, Jian, et al.
Published: (2023)
LoRA-Pro: Are Low-Rank Adapters Properly Optimized?
by: Wang, Zhengbo, et al.
Published: (2024)
by: Wang, Zhengbo, et al.
Published: (2024)
Taming Momentum: Rethinking Optimizer States Through Low-Rank Approximation
by: Wang, Zhengbo, et al.
Published: (2026)
by: Wang, Zhengbo, et al.
Published: (2026)
R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning
by: Sheng, Lijun, et al.
Published: (2025)
by: Sheng, Lijun, et al.
Published: (2025)
Exploring Vacant Classes in Label-Skewed Federated Learning
by: Guo, Kuangpu, et al.
Published: (2024)
by: Guo, Kuangpu, et al.
Published: (2024)
Adapting Vision-Language Models Without Labels: A Comprehensive Survey
by: Dong, Hao, et al.
Published: (2025)
by: Dong, Hao, et al.
Published: (2025)
Cooperative Pseudo Labeling for Unsupervised Federated Classification
by: Guo, Kuangpu, et al.
Published: (2025)
by: Guo, Kuangpu, et al.
Published: (2025)
Specialized Foundation Models Struggle to Beat Supervised Baselines
by: Xu, Zongzhe, et al.
Published: (2024)
by: Xu, Zongzhe, et al.
Published: (2024)
STAMP: Outlier-Aware Test-Time Adaptation with Stable Memory Replay
by: Yu, Yongcan, et al.
Published: (2024)
by: Yu, Yongcan, et al.
Published: (2024)
Prototypical Distillation and Debiased Tuning for Black-box Unsupervised Domain Adaptation
by: Liang, Jian, et al.
Published: (2024)
by: Liang, Jian, et al.
Published: (2024)
Strong but simple: A Baseline for Domain Generalized Dense Perception by CLIP-based Transfer Learning
by: Hümmer, Christoph, et al.
Published: (2023)
by: Hümmer, Christoph, et al.
Published: (2023)
CLIP with Generative Latent Replay: a Strong Baseline for Incremental Learning
by: Frascaroli, Emanuele, et al.
Published: (2024)
by: Frascaroli, Emanuele, et al.
Published: (2024)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
by: Lai, Zhengfeng, et al.
Published: (2023)
by: Lai, Zhengfeng, et al.
Published: (2023)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Frustratingly Easy Feature Reconstruction for Out-of-Distribution Detection
by: Wang, Yingsheng, et al.
Published: (2025)
by: Wang, Yingsheng, et al.
Published: (2025)
Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
by: Wu, Junfei, et al.
Published: (2024)
by: Wu, Junfei, et al.
Published: (2024)
Is Less More? Exploring Token Condensation as Training-free Test-time Adaptation
by: Wang, Zixin, et al.
Published: (2024)
by: Wang, Zixin, et al.
Published: (2024)
Learning Fourier shapes to probe the geometric world of deep neural networks
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging
by: Guo, Kuangpu, et al.
Published: (2025)
by: Guo, Kuangpu, et al.
Published: (2025)
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
by: Wu, Junfei, et al.
Published: (2025)
by: Wu, Junfei, et al.
Published: (2025)
ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models
by: Yin, Hao, et al.
Published: (2025)
by: Yin, Hao, et al.
Published: (2025)
Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference
by: Yin, Hao, et al.
Published: (2025)
by: Yin, Hao, et al.
Published: (2025)
LEAD: Learning Decomposition for Source-free Universal Domain Adaptation
by: Qu, Sanqing, et al.
Published: (2024)
by: Qu, Sanqing, et al.
Published: (2024)
COSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time Adaptation
by: Huang, Fanding, et al.
Published: (2025)
by: Huang, Fanding, et al.
Published: (2025)
SeMoBridge: Semantic Modality Bridge for Efficient Few-Shot Adaptation of CLIP
by: Timmermann, Christoph, et al.
Published: (2025)
by: Timmermann, Christoph, et al.
Published: (2025)
DiffCLIP: Differential Attention Meets CLIP
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
Sample Correlation for Fingerprinting Deep Face Recognition
by: Guan, Jiyang, et al.
Published: (2024)
by: Guan, Jiyang, et al.
Published: (2024)
Chatting with Images for Introspective Visual Thinking
by: Wu, Junfei, et al.
Published: (2026)
by: Wu, Junfei, et al.
Published: (2026)
MetaFormer Baselines for Vision
by: Yu, Weihao, et al.
Published: (2022)
by: Yu, Weihao, et al.
Published: (2022)
Self-Training with Dynamic Weighting for Robust Gradual Domain Adaptation
by: Wang, Zixi, et al.
Published: (2025)
by: Wang, Zixi, et al.
Published: (2025)
One Head Eight Arms: Block Matrix based Low Rank Adaptation for CLIP-based Few-Shot Learning
by: Zhou, Chunpeng, et al.
Published: (2025)
by: Zhou, Chunpeng, et al.
Published: (2025)
Explicit Uncertainty Modeling for Active CLIP Adaptation with Dual Prompt Tuning
by: Wang, Qian-Wei, et al.
Published: (2026)
by: Wang, Qian-Wei, et al.
Published: (2026)
No Captions, No Problem: Captionless 3D-CLIP Alignment with Hard Negatives via CLIP Knowledge and LLMs
by: Sbrolli, Cristian, et al.
Published: (2024)
by: Sbrolli, Cristian, et al.
Published: (2024)
CountCLIP -- [Re] Teaching CLIP to Count to Ten
by: Mestha, Harshvardhan, et al.
Published: (2024)
by: Mestha, Harshvardhan, et al.
Published: (2024)
IDEA: Image Description Enhanced CLIP-Adapter
by: Ye, Zhipeng, et al.
Published: (2025)
by: Ye, Zhipeng, et al.
Published: (2025)
Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
by: Mayilvahanan, Prasanna, et al.
Published: (2023)
by: Mayilvahanan, Prasanna, et al.
Published: (2023)
Similar Items
-
Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language Models
by: Wang, Zhengbo, et al.
Published: (2024) -
Realistic Unsupervised CLIP Fine-tuning with Universal Entropy Optimization
by: Liang, Jian, et al.
Published: (2023) -
Towards Compatible Fine-tuning for Vision-Language Model Updates
by: Wang, Zhengbo, et al.
Published: (2024) -
The Illusion of Progress? A Critical Look at Test-Time Adaptation for Vision-Language Models
by: Sheng, Lijun, et al.
Published: (2025) -
A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts
by: Liang, Jian, et al.
Published: (2023)