Logits DeConfusion with CLIP for Few-Shot Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Shuo, Liu, Fang, Hao, Zehua, Wang, Xinyi, Li, Lingling, Liu, Xu, Chen, Puhua, Ma, Wenping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913796504682496
author Li, Shuo
Liu, Fang
Hao, Zehua
Wang, Xinyi
Li, Lingling
Liu, Xu
Chen, Puhua
Ma, Wenping
author_facet Li, Shuo
Liu, Fang
Hao, Zehua
Wang, Xinyi
Li, Lingling
Liu, Xu
Chen, Puhua
Ma, Wenping
contents With its powerful visual-language alignment capability, CLIP performs well in zero-shot and few-shot learning tasks. However, we found in experiments that CLIP's logits suffer from serious inter-class confusion problems in downstream tasks, and the ambiguity between categories seriously affects the accuracy. To address this challenge, we propose a novel method called Logits DeConfusion, which effectively learns and eliminates inter-class confusion in logits by combining our Multi-level Adapter Fusion (MAF) module with our Inter-Class Deconfusion (ICD) module. Our MAF extracts features from different levels and fuses them uniformly to enhance feature representation. Our ICD learnably eliminates inter-class confusion in logits with a residual structure. Experimental results show that our method can significantly improve the classification performance and alleviate the inter-class confusion problem. The code is available at https://github.com/LiShuo1001/LDC.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12104
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Logits DeConfusion with CLIP for Few-Shot Learning
Li, Shuo
Liu, Fang
Hao, Zehua
Wang, Xinyi
Li, Lingling
Liu, Xu
Chen, Puhua
Ma, Wenping
Computer Vision and Pattern Recognition
With its powerful visual-language alignment capability, CLIP performs well in zero-shot and few-shot learning tasks. However, we found in experiments that CLIP's logits suffer from serious inter-class confusion problems in downstream tasks, and the ambiguity between categories seriously affects the accuracy. To address this challenge, we propose a novel method called Logits DeConfusion, which effectively learns and eliminates inter-class confusion in logits by combining our Multi-level Adapter Fusion (MAF) module with our Inter-Class Deconfusion (ICD) module. Our MAF extracts features from different levels and fuses them uniformly to enhance feature representation. Our ICD learnably eliminates inter-class confusion in logits with a residual structure. Experimental results show that our method can significantly improve the classification performance and alleviate the inter-class confusion problem. The code is available at https://github.com/LiShuo1001/LDC.
title Logits DeConfusion with CLIP for Few-Shot Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.12104