Evidence Packing for Cross-Domain Image Deepfake Detection with LVLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yuxin, Wang, Fei, Li, Kun, Nie, Yiqi, Chen, Junjie, Duan, Zhangling, Jia, Zhaohong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908897689731072
author Liu, Yuxin
Wang, Fei
Li, Kun
Nie, Yiqi
Chen, Junjie
Duan, Zhangling
Jia, Zhaohong
author_facet Liu, Yuxin
Wang, Fei
Li, Kun
Nie, Yiqi
Chen, Junjie
Duan, Zhangling
Jia, Zhaohong
contents Image Deepfake Detection (IDD) separates manipulated images from authentic ones by spotting artifacts of synthesis or tampering. Although large vision-language models (LVLMs) offer strong image understanding, adapting them to IDD often demands costly fine-tuning and generalizes poorly to diverse, evolving manipulations. We propose the Semantic Consistent Evidence Pack (SCEP), a training-free LVLM framework that replaces whole-image inference with evidence-driven reasoning. SCEP mines a compact set of suspicious patch tokens that best reveal manipulation cues. It uses the vision encoder's CLS token as a global reference, clusters patch features into coherent groups, and scores patches with a fused metric combining CLS-guided semantic mismatch with frequency-and noise-based anomalies. To cover dispersed traces and avoid redundancy, SCEP samples a few high-confidence patches per cluster and applies grid-based NMS, producing an evidence pack that conditions a frozen LVLM for prediction. Experiments on diverse benchmarks show SCEP outperforms strong baselines without LVLM fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17761
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evidence Packing for Cross-Domain Image Deepfake Detection with LVLMs
Liu, Yuxin
Wang, Fei
Li, Kun
Nie, Yiqi
Chen, Junjie
Duan, Zhangling
Jia, Zhaohong
Computer Vision and Pattern Recognition
Image Deepfake Detection (IDD) separates manipulated images from authentic ones by spotting artifacts of synthesis or tampering. Although large vision-language models (LVLMs) offer strong image understanding, adapting them to IDD often demands costly fine-tuning and generalizes poorly to diverse, evolving manipulations. We propose the Semantic Consistent Evidence Pack (SCEP), a training-free LVLM framework that replaces whole-image inference with evidence-driven reasoning. SCEP mines a compact set of suspicious patch tokens that best reveal manipulation cues. It uses the vision encoder's CLS token as a global reference, clusters patch features into coherent groups, and scores patches with a fused metric combining CLS-guided semantic mismatch with frequency-and noise-based anomalies. To cover dispersed traces and avoid redundancy, SCEP samples a few high-confidence patches per cluster and applies grid-based NMS, producing an evidence pack that conditions a frozen LVLM for prediction. Experiments on diverse benchmarks show SCEP outperforms strong baselines without LVLM fine-tuning.
title Evidence Packing for Cross-Domain Image Deepfake Detection with LVLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.17761