Provably Improving Generalization of Few-Shot Models with Synthetic Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Lan-Cuong, Nguyen-Tri, Quan, Khanh, Bang Tran, Le, Dung D., Tran-Thanh, Long, Than, Khoat
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911020896747520
author Nguyen, Lan-Cuong
Nguyen-Tri, Quan
Khanh, Bang Tran
Le, Dung D.
Tran-Thanh, Long
Than, Khoat
author_facet Nguyen, Lan-Cuong
Nguyen-Tri, Quan
Khanh, Bang Tran
Le, Dung D.
Tran-Thanh, Long
Than, Khoat
contents Few-shot image classification remains challenging due to the scarcity of labeled training examples. Augmenting them with synthetic data has emerged as a promising way to alleviate this issue, but models trained on synthetic samples often face performance degradation due to the inherent gap between real and synthetic distributions. To address this limitation, we develop a theoretical framework that quantifies the impact of such distribution discrepancies on supervised learning, specifically in the context of image classification. More importantly, our framework suggests practical ways to generate good synthetic samples and to train a predictor with high generalization ability. Building upon this framework, we propose a novel theoretical-based algorithm that integrates prototype learning to optimize both data partitioning and model training, effectively bridging the gap between real few-shot data and synthetic data. Extensive experiments results show that our approach demonstrates superior performance compared to state-of-the-art methods, outperforming them across multiple datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24190
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Provably Improving Generalization of Few-Shot Models with Synthetic Data
Nguyen, Lan-Cuong
Nguyen-Tri, Quan
Khanh, Bang Tran
Le, Dung D.
Tran-Thanh, Long
Than, Khoat
Machine Learning
Computer Vision and Pattern Recognition
Few-shot image classification remains challenging due to the scarcity of labeled training examples. Augmenting them with synthetic data has emerged as a promising way to alleviate this issue, but models trained on synthetic samples often face performance degradation due to the inherent gap between real and synthetic distributions. To address this limitation, we develop a theoretical framework that quantifies the impact of such distribution discrepancies on supervised learning, specifically in the context of image classification. More importantly, our framework suggests practical ways to generate good synthetic samples and to train a predictor with high generalization ability. Building upon this framework, we propose a novel theoretical-based algorithm that integrates prototype learning to optimize both data partitioning and model training, effectively bridging the gap between real few-shot data and synthetic data. Extensive experiments results show that our approach demonstrates superior performance compared to state-of-the-art methods, outperforming them across multiple datasets.
title Provably Improving Generalization of Few-Shot Models with Synthetic Data
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.24190