Memory-Anchored Multimodal Reasoning for Explainable Video Forensics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Chen, Li, Runze, Zhang, Zejun, Zhao, Pukun, Zhou, Fanqing, Wang, Longxiang, Huang, Haojian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912580303323136
author Chen, Chen
Li, Runze
Zhang, Zejun
Zhao, Pukun
Zhou, Fanqing
Wang, Longxiang
Huang, Haojian
author_facet Chen, Chen
Li, Runze
Zhang, Zejun
Zhao, Pukun
Zhou, Fanqing
Wang, Longxiang
Huang, Haojian
contents We address multimodal deepfake detection requiring both robustness and interpretability by proposing FakeHunter, a unified framework that combines memory guided retrieval, a structured Observation-Thought-Action reasoning loop, and adaptive forensic tool invocation. Visual representations from a Contrastive Language-Image Pretraining (CLIP) model and audio representations from a Contrastive Language-Audio Pretraining (CLAP) model retrieve semantically aligned authentic exemplars from a large scale memory, providing contextual anchors that guide iterative localization and explanation of suspected manipulations. Under low internal confidence the framework selectively triggers fine grained analyses such as spatial region zoom and mel spectrogram inspection to gather discriminative evidence instead of relying on opaque marginal scores. We also release X-AVFake, a comprehensive audio visual forgery benchmark with fine grained annotations of manipulation type, affected region or entity, reasoning category, and explanatory justification, designed to stress contextual grounding and explanation fidelity. Extensive experiments show that FakeHunter surpasses strong multimodal baselines, and ablation studies confirm that both contextual retrieval and selective tool activation are indispensable for improved robustness and explanatory precision.
format Preprint
id arxiv_https___arxiv_org_abs_2508_14581
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
Chen, Chen
Li, Runze
Zhang, Zejun
Zhao, Pukun
Zhou, Fanqing
Wang, Longxiang
Huang, Haojian
Multimedia
Image and Video Processing
We address multimodal deepfake detection requiring both robustness and interpretability by proposing FakeHunter, a unified framework that combines memory guided retrieval, a structured Observation-Thought-Action reasoning loop, and adaptive forensic tool invocation. Visual representations from a Contrastive Language-Image Pretraining (CLIP) model and audio representations from a Contrastive Language-Audio Pretraining (CLAP) model retrieve semantically aligned authentic exemplars from a large scale memory, providing contextual anchors that guide iterative localization and explanation of suspected manipulations. Under low internal confidence the framework selectively triggers fine grained analyses such as spatial region zoom and mel spectrogram inspection to gather discriminative evidence instead of relying on opaque marginal scores. We also release X-AVFake, a comprehensive audio visual forgery benchmark with fine grained annotations of manipulation type, affected region or entity, reasoning category, and explanatory justification, designed to stress contextual grounding and explanation fidelity. Extensive experiments show that FakeHunter surpasses strong multimodal baselines, and ablation studies confirm that both contextual retrieval and selective tool activation are indispensable for improved robustness and explanatory precision.
title Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
topic Multimedia
Image and Video Processing
url https://arxiv.org/abs/2508.14581