Logits DeConfusion with CLIP for Few-Shot Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913796504682496 |
|---|---|
| author | Li, Shuo Liu, Fang Hao, Zehua Wang, Xinyi Li, Lingling Liu, Xu Chen, Puhua Ma, Wenping |
| author_facet | Li, Shuo Liu, Fang Hao, Zehua Wang, Xinyi Li, Lingling Liu, Xu Chen, Puhua Ma, Wenping |
| contents | With its powerful visual-language alignment capability, CLIP performs well in zero-shot and few-shot learning tasks. However, we found in experiments that CLIP's logits suffer from serious inter-class confusion problems in downstream tasks, and the ambiguity between categories seriously affects the accuracy. To address this challenge, we propose a novel method called Logits DeConfusion, which effectively learns and eliminates inter-class confusion in logits by combining our Multi-level Adapter Fusion (MAF) module with our Inter-Class Deconfusion (ICD) module. Our MAF extracts features from different levels and fuses them uniformly to enhance feature representation. Our ICD learnably eliminates inter-class confusion in logits with a residual structure. Experimental results show that our method can significantly improve the classification performance and alleviate the inter-class confusion problem. The code is available at https://github.com/LiShuo1001/LDC. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_12104 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Logits DeConfusion with CLIP for Few-Shot Learning Li, Shuo Liu, Fang Hao, Zehua Wang, Xinyi Li, Lingling Liu, Xu Chen, Puhua Ma, Wenping Computer Vision and Pattern Recognition With its powerful visual-language alignment capability, CLIP performs well in zero-shot and few-shot learning tasks. However, we found in experiments that CLIP's logits suffer from serious inter-class confusion problems in downstream tasks, and the ambiguity between categories seriously affects the accuracy. To address this challenge, we propose a novel method called Logits DeConfusion, which effectively learns and eliminates inter-class confusion in logits by combining our Multi-level Adapter Fusion (MAF) module with our Inter-Class Deconfusion (ICD) module. Our MAF extracts features from different levels and fuses them uniformly to enhance feature representation. Our ICD learnably eliminates inter-class confusion in logits with a residual structure. Experimental results show that our method can significantly improve the classification performance and alleviate the inter-class confusion problem. The code is available at https://github.com/LiShuo1001/LDC. |
| title | Logits DeConfusion with CLIP for Few-Shot Learning |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2504.12104 |