ZeroDiff: Solidified Visual-Semantic Correlation in Zero-Shot Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ye, Zihan, Gowda, Shreyank N., Huang, Xiaowei, Xu, Haotian, Jin, Yaochu, Huang, Kaizhu, Jin, Xiaobo
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929708386484224
author Ye, Zihan
Gowda, Shreyank N.
Huang, Xiaowei
Xu, Haotian
Jin, Yaochu
Huang, Kaizhu
Jin, Xiaobo
author_facet Ye, Zihan
Gowda, Shreyank N.
Huang, Xiaowei
Xu, Haotian
Jin, Yaochu
Huang, Kaizhu
Jin, Xiaobo
contents Zero-shot Learning (ZSL) aims to enable classifiers to identify unseen classes. This is typically achieved by generating visual features for unseen classes based on learned visual-semantic correlations from seen classes. However, most current generative approaches heavily rely on having a sufficient number of samples from seen classes. Our study reveals that a scarcity of seen class samples results in a marked decrease in performance across many generative ZSL techniques. We argue, quantify, and empirically demonstrate that this decline is largely attributable to spurious visual-semantic correlations. To address this issue, we introduce ZeroDiff, an innovative generative framework for ZSL that incorporates diffusion mechanisms and contrastive representations to enhance visual-semantic correlations. ZeroDiff comprises three key components: (1) Diffusion augmentation, which naturally transforms limited data into an expanded set of noised data to mitigate generative model overfitting; (2) Supervised-contrastive (SC)-based representations that dynamically characterize each limited sample to support visual feature generation; and (3) Multiple feature discriminators employing a Wasserstein-distance-based mutual learning approach, evaluating generated features from various perspectives, including pre-defined semantics, SC-based representations, and the diffusion process. Extensive experiments on three popular ZSL benchmarks demonstrate that ZeroDiff not only achieves significant improvements over existing ZSL methods but also maintains robust performance even with scarce training data. Our codes are available at https://github.com/FouriYe/ZeroDiff_ICLR25.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02929
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ZeroDiff: Solidified Visual-Semantic Correlation in Zero-Shot Learning
Ye, Zihan
Gowda, Shreyank N.
Huang, Xiaowei
Xu, Haotian
Jin, Yaochu
Huang, Kaizhu
Jin, Xiaobo
Computer Vision and Pattern Recognition
Machine Learning
Zero-shot Learning (ZSL) aims to enable classifiers to identify unseen classes. This is typically achieved by generating visual features for unseen classes based on learned visual-semantic correlations from seen classes. However, most current generative approaches heavily rely on having a sufficient number of samples from seen classes. Our study reveals that a scarcity of seen class samples results in a marked decrease in performance across many generative ZSL techniques. We argue, quantify, and empirically demonstrate that this decline is largely attributable to spurious visual-semantic correlations. To address this issue, we introduce ZeroDiff, an innovative generative framework for ZSL that incorporates diffusion mechanisms and contrastive representations to enhance visual-semantic correlations. ZeroDiff comprises three key components: (1) Diffusion augmentation, which naturally transforms limited data into an expanded set of noised data to mitigate generative model overfitting; (2) Supervised-contrastive (SC)-based representations that dynamically characterize each limited sample to support visual feature generation; and (3) Multiple feature discriminators employing a Wasserstein-distance-based mutual learning approach, evaluating generated features from various perspectives, including pre-defined semantics, SC-based representations, and the diffusion process. Extensive experiments on three popular ZSL benchmarks demonstrate that ZeroDiff not only achieves significant improvements over existing ZSL methods but also maintains robust performance even with scarce training data. Our codes are available at https://github.com/FouriYe/ZeroDiff_ICLR25.
title ZeroDiff: Solidified Visual-Semantic Correlation in Zero-Shot Learning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2406.02929