Salient Concept-Aware Generative Data Augmentation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866912653911261184 |
|---|---|
| author | Zhao, Tianchen Chen, Xuanbai Li, Zhihua Fang, Jun An, Dongsheng Xu, Xiang Tu, Zhuowen Xing, Yifan |
| author_facet | Zhao, Tianchen Chen, Xuanbai Li, Zhihua Fang, Jun An, Dongsheng Xu, Xiang Tu, Zhuowen Xing, Yifan |
| contents | Recent generative data augmentation methods conditioned on both image and text prompts struggle to balance between fidelity and diversity, as it is challenging to preserve essential image details while aligning with varied text prompts. This challenge arises because representations in the synthesis process often become entangled with non-essential input image attributes such as environmental contexts, creating conflicts with text prompts intended to modify these elements. To address this, we propose a personalized image generation framework that uses a salient concept-aware image embedding model to reduce the influence of irrelevant visual details during the synthesis process, thereby maintaining intuitive alignment between image and text inputs. By generating images that better preserve class-discriminative features with additional controlled variations, our framework effectively enhances the diversity of training datasets and thereby improves the robustness of downstream models. Our approach demonstrates superior performance across eight fine-grained vision datasets, outperforming state-of-the-art augmentation methods with averaged classification accuracy improvements by 0.73% and 6.5% under conventional and long-tail settings, respectively. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_15194 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Salient Concept-Aware Generative Data Augmentation Zhao, Tianchen Chen, Xuanbai Li, Zhihua Fang, Jun An, Dongsheng Xu, Xiang Tu, Zhuowen Xing, Yifan Computer Vision and Pattern Recognition 68T45 (Machine learning) I.2.10; I.2.6; I.4.8; I.5.1; I.5.4 Recent generative data augmentation methods conditioned on both image and text prompts struggle to balance between fidelity and diversity, as it is challenging to preserve essential image details while aligning with varied text prompts. This challenge arises because representations in the synthesis process often become entangled with non-essential input image attributes such as environmental contexts, creating conflicts with text prompts intended to modify these elements. To address this, we propose a personalized image generation framework that uses a salient concept-aware image embedding model to reduce the influence of irrelevant visual details during the synthesis process, thereby maintaining intuitive alignment between image and text inputs. By generating images that better preserve class-discriminative features with additional controlled variations, our framework effectively enhances the diversity of training datasets and thereby improves the robustness of downstream models. Our approach demonstrates superior performance across eight fine-grained vision datasets, outperforming state-of-the-art augmentation methods with averaged classification accuracy improvements by 0.73% and 6.5% under conventional and long-tail settings, respectively. |
| title | Salient Concept-Aware Generative Data Augmentation |
| topic | Computer Vision and Pattern Recognition 68T45 (Machine learning) I.2.10; I.2.6; I.4.8; I.5.1; I.5.4 |
| url | https://arxiv.org/abs/2510.15194 |