Towards Explainable Fake Image Detection with Multi-Modal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ji, Yikun, Hong, Yan, Zhan, Jiahui, Chen, Haoxing, lan, jun, Zhu, Huijia, Wang, Weiqiang, Zhang, Liqing, Zhang, Jianfu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915604697448448
author Ji, Yikun
Hong, Yan
Zhan, Jiahui
Chen, Haoxing
lan, jun
Zhu, Huijia
Wang, Weiqiang
Zhang, Liqing
Zhang, Jianfu
author_facet Ji, Yikun
Hong, Yan
Zhan, Jiahui
Chen, Haoxing
lan, jun
Zhu, Huijia
Wang, Weiqiang
Zhang, Liqing
Zhang, Jianfu
contents Progress in image generation raises significant public security concerns. We argue that fake image detection should not operate as a "black box". Instead, an ideal approach must ensure both strong generalization and transparency. Recent progress in Multi-modal Large Language Models (MLLMs) offers new opportunities for reasoning-based AI-generated image detection. In this work, we evaluate the capabilities of MLLMs in comparison to traditional detection methods and human evaluators, highlighting their strengths and limitations. Furthermore, we design six distinct prompts and propose a framework that integrates these prompts to develop a more robust, explainable, and reasoning-driven detection system. The code is available at https://github.com/Gennadiyev/mllm-defake.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14245
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
Ji, Yikun
Hong, Yan
Zhan, Jiahui
Chen, Haoxing
lan, jun
Zhu, Huijia
Wang, Weiqiang
Zhang, Liqing
Zhang, Jianfu
Computer Vision and Pattern Recognition
Computation and Language
I.2.7; I.2.10
Progress in image generation raises significant public security concerns. We argue that fake image detection should not operate as a "black box". Instead, an ideal approach must ensure both strong generalization and transparency. Recent progress in Multi-modal Large Language Models (MLLMs) offers new opportunities for reasoning-based AI-generated image detection. In this work, we evaluate the capabilities of MLLMs in comparison to traditional detection methods and human evaluators, highlighting their strengths and limitations. Furthermore, we design six distinct prompts and propose a framework that integrates these prompts to develop a more robust, explainable, and reasoning-driven detection system. The code is available at https://github.com/Gennadiyev/mllm-defake.
title Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
I.2.7; I.2.10
url https://arxiv.org/abs/2504.14245