IDEA: Image Description Enhanced CLIP-Adapter
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ye, Zhipeng, Jiang, Feng, Wang, Qiufeng, Huang, Kaizhu, Huang, Jiaqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification
von: Tan, Zhaorui, et al.
Veröffentlicht: (2024)
von: Tan, Zhaorui, et al.
Veröffentlicht: (2024)
Rethinking Multi-domain Generalization with A General Learning Objective
von: Tan, Zhaorui, et al.
Veröffentlicht: (2024)
von: Tan, Zhaorui, et al.
Veröffentlicht: (2024)
A generalizable framework for low-rank tensor completion with numerical priors
von: Yuan, Shiran, et al.
Veröffentlicht: (2023)
von: Yuan, Shiran, et al.
Veröffentlicht: (2023)
Divide-and-Conquer Inference for Large-Scale Visual Recognition with Multimodal Large Language Models
von: Ye, Zhipeng, et al.
Veröffentlicht: (2026)
von: Ye, Zhipeng, et al.
Veröffentlicht: (2026)
RankCLIP: Ranking-Consistent Language-Image Pretraining
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
KernelDNA: Dynamic Kernel Sharing via Decoupled Naive Adapters
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
COSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time Adaptation
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
The Demon is in Ambiguity: Revisiting Situation Recognition with Single Positive Multi-Label Learning
von: Lin, Yiming, et al.
Veröffentlicht: (2025)
von: Lin, Yiming, et al.
Veröffentlicht: (2025)
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
von: Lin, Feng, et al.
Veröffentlicht: (2025)
von: Lin, Feng, et al.
Veröffentlicht: (2025)
ATLAS: Adapter-Based Multi-Modal Continual Learning with a Two-Stage Learning Strategy
von: Li, Hong, et al.
Veröffentlicht: (2024)
von: Li, Hong, et al.
Veröffentlicht: (2024)
RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
von: Lin, Xiang, et al.
Veröffentlicht: (2025)
von: Lin, Xiang, et al.
Veröffentlicht: (2025)
Unraveling Batch Normalization for Realistic Test-Time Adaptation
von: Su, Zixian, et al.
Veröffentlicht: (2023)
von: Su, Zixian, et al.
Veröffentlicht: (2023)
DiffCLIP: Differential Attention Meets CLIP
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
von: Crabbé, Jonathan, et al.
Veröffentlicht: (2023)
von: Crabbé, Jonathan, et al.
Veröffentlicht: (2023)
AdapterTune: Zero-Initialized Low-Rank Adapters for Frozen Vision Transformers
von: Khazem, Salim
Veröffentlicht: (2026)
von: Khazem, Salim
Veröffentlicht: (2026)
CountCLIP -- [Re] Teaching CLIP to Count to Ten
von: Mestha, Harshvardhan, et al.
Veröffentlicht: (2024)
von: Mestha, Harshvardhan, et al.
Veröffentlicht: (2024)
Hyperbolic Structured Classification for Robust Single Positive Multi-label Learning
von: Lin, Yiming, et al.
Veröffentlicht: (2025)
von: Lin, Yiming, et al.
Veröffentlicht: (2025)
Rethinking Information Loss in Medical Image Segmentation with Various-sized Targets
von: Liu, Tianyi, et al.
Veröffentlicht: (2024)
von: Liu, Tianyi, et al.
Veröffentlicht: (2024)
Synergy and Diversity in CLIP: Enhancing Performance Through Adaptive Backbone Ensembling
von: Rodriguez-Opazo, Cristian, et al.
Veröffentlicht: (2024)
von: Rodriguez-Opazo, Cristian, et al.
Veröffentlicht: (2024)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
von: Lai, Zhengfeng, et al.
Veröffentlicht: (2023)
von: Lai, Zhengfeng, et al.
Veröffentlicht: (2023)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
GoldiCLIP: The Goldilocks Approach for Balancing Explicit Supervision for Language-Image Pretraining
von: Mohan, Deen Dayal, et al.
Veröffentlicht: (2026)
von: Mohan, Deen Dayal, et al.
Veröffentlicht: (2026)
A comprehensive survey of oracle character recognition: challenges, benchmarks, and beyond
von: Li, Jing, et al.
Veröffentlicht: (2024)
von: Li, Jing, et al.
Veröffentlicht: (2024)
Semantic Relation-Enhanced CLIP Adapter for Domain Adaptive Zero-Shot Learning
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
Dynamic Integration of Task-Specific Adapters for Class Incremental Learning
von: Li, Jiashuo, et al.
Veröffentlicht: (2024)
von: Li, Jiashuo, et al.
Veröffentlicht: (2024)
AutoGeo: Automating Geometric Image Dataset Creation for Enhanced Geometry Understanding
von: Huang, Zihan, et al.
Veröffentlicht: (2024)
von: Huang, Zihan, et al.
Veröffentlicht: (2024)
Reproducibility Study of CDUL: CLIP-Driven Unsupervised Learning for Multi-Label Image Classification
von: Shah, Manan, et al.
Veröffentlicht: (2024)
von: Shah, Manan, et al.
Veröffentlicht: (2024)
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts
von: An, Bang, et al.
Veröffentlicht: (2023)
von: An, Bang, et al.
Veröffentlicht: (2023)
Training-Only Heterogeneous Image-Patch-Text Graph Supervision for Advancing Few-Shot Learning Adapters
von: Mohammad, Mohammed Rahman Sherif Khan, et al.
Veröffentlicht: (2026)
von: Mohammad, Mohammed Rahman Sherif Khan, et al.
Veröffentlicht: (2026)
Enhancing CLIP Conceptual Embedding through Knowledge Distillation
von: Kao, Kuei-Chun
Veröffentlicht: (2024)
von: Kao, Kuei-Chun
Veröffentlicht: (2024)
CLIP Can Understand Depth
von: Kim, Sohee, et al.
Veröffentlicht: (2024)
von: Kim, Sohee, et al.
Veröffentlicht: (2024)
DAM: Dynamic Adapter Merging for Continual Video QA Learning
von: Cheng, Feng, et al.
Veröffentlicht: (2024)
von: Cheng, Feng, et al.
Veröffentlicht: (2024)
Multi-Modal Adapter for Vision-Language Models
von: Seputis, Dominykas, et al.
Veröffentlicht: (2024)
von: Seputis, Dominykas, et al.
Veröffentlicht: (2024)
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
von: Huang, Hanxun, et al.
Veröffentlicht: (2025)
von: Huang, Hanxun, et al.
Veröffentlicht: (2025)
Captured by Captions: On Memorization and its Mitigation in CLIP Models
von: Wang, Wenhao, et al.
Veröffentlicht: (2025)
von: Wang, Wenhao, et al.
Veröffentlicht: (2025)
CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP
von: Yang, Tianyu, et al.
Veröffentlicht: (2024)
von: Yang, Tianyu, et al.
Veröffentlicht: (2024)
Interpreting CLIP with Hierarchical Sparse Autoencoders
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2025)
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2025)
ECOR: Explainable CLIP for Object Recognition
von: Rasekh, Ali, et al.
Veröffentlicht: (2024)
von: Rasekh, Ali, et al.
Veröffentlicht: (2024)
A Hard-to-Beat Baseline for Training-free CLIP-based Adaptation
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
Continual Diffusion with STAMINA: STack-And-Mask INcremental Adapters
von: Smith, James Seale, et al.
Veröffentlicht: (2023)
von: Smith, James Seale, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification
von: Tan, Zhaorui, et al.
Veröffentlicht: (2024) -
Rethinking Multi-domain Generalization with A General Learning Objective
von: Tan, Zhaorui, et al.
Veröffentlicht: (2024) -
A generalizable framework for low-rank tensor completion with numerical priors
von: Yuan, Shiran, et al.
Veröffentlicht: (2023) -
Divide-and-Conquer Inference for Large-Scale Visual Recognition with Multimodal Large Language Models
von: Ye, Zhipeng, et al.
Veröffentlicht: (2026) -
RankCLIP: Ranking-Consistent Language-Image Pretraining
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)