GAMMA: Generalizable Alignment via Multi-task and Manipulation-Augmented Training for AI-Generated Image Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yan, Haozhen, Hong, Yan, Lang, Suning, Zhan, Jiahui, Ji, Yikun, Gao, Yujie, Zhu, Huijia, Lan, Jun, Zhang, Jianfu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915748813733888
author Yan, Haozhen
Hong, Yan
Lang, Suning
Zhan, Jiahui
Ji, Yikun
Gao, Yujie
Zhu, Huijia
Lan, Jun
Zhang, Jianfu
author_facet Yan, Haozhen
Hong, Yan
Lang, Suning
Zhan, Jiahui
Ji, Yikun
Gao, Yujie
Zhu, Huijia
Lan, Jun
Zhang, Jianfu
contents With generative models becoming increasingly sophisticated and diverse, detecting AI-generated images has become increasingly challenging. While existing AI-genereted Image detectors achieve promising performance on in-distribution generated images, their generalization to unseen generative models remains limited. This limitation is largely attributed to their reliance on generation-specific artifacts, such as stylistic priors and compression patterns. To address these limitations, we propose GAMMA, a novel training framework designed to reduce domain bias and enhance semantic alignment. GAMMA introduces diverse manipulation strategies, such as inpainting-based manipulation and semantics-preserving perturbations, to ensure consistency between manipulated and authentic content. We employ multi-task supervision with dual segmentation heads and a classification head, enabling pixel-level source attribution across diverse generative domains. In addition, a reverse cross-attention mechanism is introduced to allow the segmentation heads to guide and correct biased representations in the classification branch. Our method achieves state-of-the-art generalization performance on the GenImage benchmark, imporving accuracy by 5.8%, but also maintains strong robustness on newly released generative model such as GPT-4o.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10250
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GAMMA: Generalizable Alignment via Multi-task and Manipulation-Augmented Training for AI-Generated Image Detection
Yan, Haozhen
Hong, Yan
Lang, Suning
Zhan, Jiahui
Ji, Yikun
Gao, Yujie
Zhu, Huijia
Lan, Jun
Zhang, Jianfu
Computer Vision and Pattern Recognition
With generative models becoming increasingly sophisticated and diverse, detecting AI-generated images has become increasingly challenging. While existing AI-genereted Image detectors achieve promising performance on in-distribution generated images, their generalization to unseen generative models remains limited. This limitation is largely attributed to their reliance on generation-specific artifacts, such as stylistic priors and compression patterns. To address these limitations, we propose GAMMA, a novel training framework designed to reduce domain bias and enhance semantic alignment. GAMMA introduces diverse manipulation strategies, such as inpainting-based manipulation and semantics-preserving perturbations, to ensure consistency between manipulated and authentic content. We employ multi-task supervision with dual segmentation heads and a classification head, enabling pixel-level source attribution across diverse generative domains. In addition, a reverse cross-attention mechanism is introduced to allow the segmentation heads to guide and correct biased representations in the classification branch. Our method achieves state-of-the-art generalization performance on the GenImage benchmark, imporving accuracy by 5.8%, but also maintains strong robustness on newly released generative model such as GPT-4o.
title GAMMA: Generalizable Alignment via Multi-task and Manipulation-Augmented Training for AI-Generated Image Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.10250