Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dong, Yixuan, Su, Fang-Yi, Chiang, Jung-Hsien
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:https://arxiv.org/abs/2505.11813
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908368796385280
author Dong, Yixuan
Su, Fang-Yi
Chiang, Jung-Hsien
author_facet Dong, Yixuan
Su, Fang-Yi
Chiang, Jung-Hsien
contents Data augmentation for domain-specific image classification tasks often struggles to simultaneously address diversity, faithfulness, and label clarity of generated data, leading to suboptimal performance in downstream tasks. While existing generative diffusion model-based methods aim to enhance augmentation, they fail to cohesively tackle these three critical aspects and often overlook intrinsic challenges of diffusion models, such as sensitivity to model characteristics and stochasticity under strong transformations. In this paper, we propose a novel framework that explicitly integrates diversity, faithfulness, and label clarity into the augmentation process. Our approach employs saliency-guided mixing and a fine-tuned diffusion model to preserve foreground semantics, enrich background diversity, and ensure label consistency, while mitigating diffusion model limitations. Extensive experiments across fine-grained, long-tail, few-shot, and background robustness tasks demonstrate our method's superior performance over state-of-the-art approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11813
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SGD-Mix: Enhancing Domain-Specific Image Classification with Label-Preserving Data Augmentation
Dong, Yixuan
Su, Fang-Yi
Chiang, Jung-Hsien
Computer Vision and Pattern Recognition
Artificial Intelligence
Data augmentation for domain-specific image classification tasks often struggles to simultaneously address diversity, faithfulness, and label clarity of generated data, leading to suboptimal performance in downstream tasks. While existing generative diffusion model-based methods aim to enhance augmentation, they fail to cohesively tackle these three critical aspects and often overlook intrinsic challenges of diffusion models, such as sensitivity to model characteristics and stochasticity under strong transformations. In this paper, we propose a novel framework that explicitly integrates diversity, faithfulness, and label clarity into the augmentation process. Our approach employs saliency-guided mixing and a fine-tuned diffusion model to preserve foreground semantics, enrich background diversity, and ensure label consistency, while mitigating diffusion model limitations. Extensive experiments across fine-grained, long-tail, few-shot, and background robustness tasks demonstrate our method's superior performance over state-of-the-art approaches.
title SGD-Mix: Enhancing Domain-Specific Image Classification with Label-Preserving Data Augmentation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.11813