SKDU at De-Factify 4.0: Vision Transformer with Data Augmentation for AI-Generated Image Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866912290707603456 |
|---|---|
| author | Malviya, Shrikant Bhowmik, Neelanjan Katsigiannis, Stamos |
| author_facet | Malviya, Shrikant Bhowmik, Neelanjan Katsigiannis, Stamos |
| contents | The aim of this work is to explore the potential of pre-trained vision-language models, e.g. Vision Transformers (ViT), enhanced with advanced data augmentation strategies for the detection of AI-generated images. Our approach leverages a fine-tuned ViT model trained on the Defactify-4.0 dataset, which includes images generated by state-of-the-art models such as Stable Diffusion 2.1, Stable Diffusion XL, Stable Diffusion 3, DALL-E 3, and MidJourney. We employ perturbation techniques like flipping, rotation, Gaussian noise injection, and JPEG compression during training to improve model robustness and generalisation. The experimental results demonstrate that our ViT-based pipeline achieves state-of-the-art performance, significantly outperforming competing methods on both validation and test datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_18812 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SKDU at De-Factify 4.0: Vision Transformer with Data Augmentation for AI-Generated Image Detection Malviya, Shrikant Bhowmik, Neelanjan Katsigiannis, Stamos Computer Vision and Pattern Recognition The aim of this work is to explore the potential of pre-trained vision-language models, e.g. Vision Transformers (ViT), enhanced with advanced data augmentation strategies for the detection of AI-generated images. Our approach leverages a fine-tuned ViT model trained on the Defactify-4.0 dataset, which includes images generated by state-of-the-art models such as Stable Diffusion 2.1, Stable Diffusion XL, Stable Diffusion 3, DALL-E 3, and MidJourney. We employ perturbation techniques like flipping, rotation, Gaussian noise injection, and JPEG compression during training to improve model robustness and generalisation. The experimental results demonstrate that our ViT-based pipeline achieves state-of-the-art performance, significantly outperforming competing methods on both validation and test datasets. |
| title | SKDU at De-Factify 4.0: Vision Transformer with Data Augmentation for AI-Generated Image Detection |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.18812 |