Cross-Modal Few-Shot Learning with Second-Order Neural Ordinary Differential Equations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yi, Cheng, Chun-Wun, He, Junyi, He, Zhihai, Schönlieb, Carola-Bibiane, Chen, Yuyan, Aviles-Rivero, Angelica I
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909436177547264
author Zhang, Yi
Cheng, Chun-Wun
He, Junyi
He, Zhihai
Schönlieb, Carola-Bibiane
Chen, Yuyan
Aviles-Rivero, Angelica I
author_facet Zhang, Yi
Cheng, Chun-Wun
He, Junyi
He, Zhihai
Schönlieb, Carola-Bibiane
Chen, Yuyan
Aviles-Rivero, Angelica I
contents We introduce SONO, a novel method leveraging Second-Order Neural Ordinary Differential Equations (Second-Order NODEs) to enhance cross-modal few-shot learning. By employing a simple yet effective architecture consisting of a Second-Order NODEs model paired with a cross-modal classifier, SONO addresses the significant challenge of overfitting, which is common in few-shot scenarios due to limited training examples. Our second-order approach can approximate a broader class of functions, enhancing the model's expressive power and feature generalization capabilities. We initialize our cross-modal classifier with text embeddings derived from class-relevant prompts, streamlining training efficiency by avoiding the need for frequent text encoder processing. Additionally, we utilize text-based image augmentation, exploiting CLIP's robust image-text correlation to enrich training data significantly. Extensive experiments across multiple datasets demonstrate that SONO outperforms existing state-of-the-art methods in few-shot learning performance.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15813
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cross-Modal Few-Shot Learning with Second-Order Neural Ordinary Differential Equations
Zhang, Yi
Cheng, Chun-Wun
He, Junyi
He, Zhihai
Schönlieb, Carola-Bibiane
Chen, Yuyan
Aviles-Rivero, Angelica I
Computer Vision and Pattern Recognition
We introduce SONO, a novel method leveraging Second-Order Neural Ordinary Differential Equations (Second-Order NODEs) to enhance cross-modal few-shot learning. By employing a simple yet effective architecture consisting of a Second-Order NODEs model paired with a cross-modal classifier, SONO addresses the significant challenge of overfitting, which is common in few-shot scenarios due to limited training examples. Our second-order approach can approximate a broader class of functions, enhancing the model's expressive power and feature generalization capabilities. We initialize our cross-modal classifier with text embeddings derived from class-relevant prompts, streamlining training efficiency by avoiding the need for frequent text encoder processing. Additionally, we utilize text-based image augmentation, exploiting CLIP's robust image-text correlation to enrich training data significantly. Extensive experiments across multiple datasets demonstrate that SONO outperforms existing state-of-the-art methods in few-shot learning performance.
title Cross-Modal Few-Shot Learning with Second-Order Neural Ordinary Differential Equations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.15813