GenMix: Effective Data Augmentation with Generative Diffusion Model Image Editing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866912613963661312 |
|---|---|
| author | Islam, Khawar Zaheer, Muhammad Zaigham Mahmood, Arif Nandakumar, Karthik Akhtar, Naveed |
| author_facet | Islam, Khawar Zaheer, Muhammad Zaigham Mahmood, Arif Nandakumar, Karthik Akhtar, Naveed |
| contents | Data augmentation is widely used to enhance generalization in visual classification tasks. However, traditional methods struggle when source and target domains differ, as in domain adaptation, due to their inability to address domain gaps. This paper introduces GenMix, a generalizable prompt-guided generative data augmentation approach that enhances both in-domain and cross-domain image classification. Our technique leverages image editing to generate augmented images based on custom conditional prompts, designed specifically for each problem type. By blending portions of the input image with its edited generative counterpart and incorporating fractal patterns, our approach mitigates unrealistic images and label ambiguity, improving the performance and adversarial robustness of the resulting models. Efficacy of our method is established with extensive experiments on eight public datasets for general and fine-grained classification, in both in-domain and cross-domain settings. Additionally, we demonstrate performance improvements for self-supervised learning, learning with data scarcity, and adversarial robustness. As compared to the existing state-of-the-art methods, our technique achieves stronger performance across the board. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_02366 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | GenMix: Effective Data Augmentation with Generative Diffusion Model Image Editing Islam, Khawar Zaheer, Muhammad Zaigham Mahmood, Arif Nandakumar, Karthik Akhtar, Naveed Computer Vision and Pattern Recognition Data augmentation is widely used to enhance generalization in visual classification tasks. However, traditional methods struggle when source and target domains differ, as in domain adaptation, due to their inability to address domain gaps. This paper introduces GenMix, a generalizable prompt-guided generative data augmentation approach that enhances both in-domain and cross-domain image classification. Our technique leverages image editing to generate augmented images based on custom conditional prompts, designed specifically for each problem type. By blending portions of the input image with its edited generative counterpart and incorporating fractal patterns, our approach mitigates unrealistic images and label ambiguity, improving the performance and adversarial robustness of the resulting models. Efficacy of our method is established with extensive experiments on eight public datasets for general and fine-grained classification, in both in-domain and cross-domain settings. Additionally, we demonstrate performance improvements for self-supervised learning, learning with data scarcity, and adversarial robustness. As compared to the existing state-of-the-art methods, our technique achieves stronger performance across the board. |
| title | GenMix: Effective Data Augmentation with Generative Diffusion Model Image Editing |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.02366 |