Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915604697448448 |
|---|---|
| author | Ji, Yikun Hong, Yan Zhan, Jiahui Chen, Haoxing lan, jun Zhu, Huijia Wang, Weiqiang Zhang, Liqing Zhang, Jianfu |
| author_facet | Ji, Yikun Hong, Yan Zhan, Jiahui Chen, Haoxing lan, jun Zhu, Huijia Wang, Weiqiang Zhang, Liqing Zhang, Jianfu |
| contents | Progress in image generation raises significant public security concerns. We argue that fake image detection should not operate as a "black box". Instead, an ideal approach must ensure both strong generalization and transparency. Recent progress in Multi-modal Large Language Models (MLLMs) offers new opportunities for reasoning-based AI-generated image detection. In this work, we evaluate the capabilities of MLLMs in comparison to traditional detection methods and human evaluators, highlighting their strengths and limitations. Furthermore, we design six distinct prompts and propose a framework that integrates these prompts to develop a more robust, explainable, and reasoning-driven detection system. The code is available at https://github.com/Gennadiyev/mllm-defake. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_14245 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Towards Explainable Fake Image Detection with Multi-Modal Large Language Models Ji, Yikun Hong, Yan Zhan, Jiahui Chen, Haoxing lan, jun Zhu, Huijia Wang, Weiqiang Zhang, Liqing Zhang, Jianfu Computer Vision and Pattern Recognition Computation and Language I.2.7; I.2.10 Progress in image generation raises significant public security concerns. We argue that fake image detection should not operate as a "black box". Instead, an ideal approach must ensure both strong generalization and transparency. Recent progress in Multi-modal Large Language Models (MLLMs) offers new opportunities for reasoning-based AI-generated image detection. In this work, we evaluate the capabilities of MLLMs in comparison to traditional detection methods and human evaluators, highlighting their strengths and limitations. Furthermore, we design six distinct prompts and propose a framework that integrates these prompts to develop a more robust, explainable, and reasoning-driven detection system. The code is available at https://github.com/Gennadiyev/mllm-defake. |
| title | Towards Explainable Fake Image Detection with Multi-Modal Large Language Models |
| topic | Computer Vision and Pattern Recognition Computation and Language I.2.7; I.2.10 |
| url | https://arxiv.org/abs/2504.14245 |