Salient Concept-Aware Generative Data Augmentation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhao, Tianchen, Chen, Xuanbai, Li, Zhihua, Fang, Jun, An, Dongsheng, Xu, Xiang, Tu, Zhuowen, Xing, Yifan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912653911261184
author Zhao, Tianchen
Chen, Xuanbai
Li, Zhihua
Fang, Jun
An, Dongsheng
Xu, Xiang
Tu, Zhuowen
Xing, Yifan
author_facet Zhao, Tianchen
Chen, Xuanbai
Li, Zhihua
Fang, Jun
An, Dongsheng
Xu, Xiang
Tu, Zhuowen
Xing, Yifan
contents Recent generative data augmentation methods conditioned on both image and text prompts struggle to balance between fidelity and diversity, as it is challenging to preserve essential image details while aligning with varied text prompts. This challenge arises because representations in the synthesis process often become entangled with non-essential input image attributes such as environmental contexts, creating conflicts with text prompts intended to modify these elements. To address this, we propose a personalized image generation framework that uses a salient concept-aware image embedding model to reduce the influence of irrelevant visual details during the synthesis process, thereby maintaining intuitive alignment between image and text inputs. By generating images that better preserve class-discriminative features with additional controlled variations, our framework effectively enhances the diversity of training datasets and thereby improves the robustness of downstream models. Our approach demonstrates superior performance across eight fine-grained vision datasets, outperforming state-of-the-art augmentation methods with averaged classification accuracy improvements by 0.73% and 6.5% under conventional and long-tail settings, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15194
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Salient Concept-Aware Generative Data Augmentation
Zhao, Tianchen
Chen, Xuanbai
Li, Zhihua
Fang, Jun
An, Dongsheng
Xu, Xiang
Tu, Zhuowen
Xing, Yifan
Computer Vision and Pattern Recognition
68T45 (Machine learning)
I.2.10; I.2.6; I.4.8; I.5.1; I.5.4
Recent generative data augmentation methods conditioned on both image and text prompts struggle to balance between fidelity and diversity, as it is challenging to preserve essential image details while aligning with varied text prompts. This challenge arises because representations in the synthesis process often become entangled with non-essential input image attributes such as environmental contexts, creating conflicts with text prompts intended to modify these elements. To address this, we propose a personalized image generation framework that uses a salient concept-aware image embedding model to reduce the influence of irrelevant visual details during the synthesis process, thereby maintaining intuitive alignment between image and text inputs. By generating images that better preserve class-discriminative features with additional controlled variations, our framework effectively enhances the diversity of training datasets and thereby improves the robustness of downstream models. Our approach demonstrates superior performance across eight fine-grained vision datasets, outperforming state-of-the-art augmentation methods with averaged classification accuracy improvements by 0.73% and 6.5% under conventional and long-tail settings, respectively.
title Salient Concept-Aware Generative Data Augmentation
topic Computer Vision and Pattern Recognition
68T45 (Machine learning)
I.2.10; I.2.6; I.4.8; I.5.1; I.5.4
url https://arxiv.org/abs/2510.15194