X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Youngseo, Yun, Kwan, Hong, Seokhyeon, Cha, Sihun, Koo, Colette Suhjung, Noh, Junyong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918380546555904
author Kim, Youngseo
Yun, Kwan
Hong, Seokhyeon
Cha, Sihun
Koo, Colette Suhjung
Noh, Junyong
author_facet Kim, Youngseo
Yun, Kwan
Hong, Seokhyeon
Cha, Sihun
Koo, Colette Suhjung
Noh, Junyong
contents The surge of highly realistic synthetic videos produced by contemporary generative systems has significantly increased the risk of malicious use, challenging both humans and existing detectors. Against this backdrop, we take a generator-side view and observe that internal cross-attention mechanisms in these models encode fine-grained speech-motion alignment, offering useful correspondence cues for forgery detection. Building on this insight, we propose X-AVDT, a robust and generalizable deepfake detector that probes generator-internal audio-visual signals accessed via DDIM inversion to expose these cues. X-AVDT extracts two complementary signals: (i) a video composite capturing inversion-induced discrepancies, and (ii) an audio-visual cross-attention feature reflecting modality alignment enforced during generation. To enable faithful cross-generator evaluation, we further introduce MMDF, a new multimodal deepfake dataset spanning diverse manipulation types and rapidly evolving synthesis paradigms, including GANs, diffusion, and flow-matching. Extensive experiments demonstrate that X-AVDT achieves leading performance on MMDF and generalizes strongly to external benchmarks and unseen generators, outperforming existing methods with accuracy improved by 13.1%. Our findings highlight the importance of leveraging internal audio-visual consistency cues for robustness to future generators in deepfake detection.
format Preprint
id arxiv_https___arxiv_org_abs_2603_08483
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection
Kim, Youngseo
Yun, Kwan
Hong, Seokhyeon
Cha, Sihun
Koo, Colette Suhjung
Noh, Junyong
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
The surge of highly realistic synthetic videos produced by contemporary generative systems has significantly increased the risk of malicious use, challenging both humans and existing detectors. Against this backdrop, we take a generator-side view and observe that internal cross-attention mechanisms in these models encode fine-grained speech-motion alignment, offering useful correspondence cues for forgery detection. Building on this insight, we propose X-AVDT, a robust and generalizable deepfake detector that probes generator-internal audio-visual signals accessed via DDIM inversion to expose these cues. X-AVDT extracts two complementary signals: (i) a video composite capturing inversion-induced discrepancies, and (ii) an audio-visual cross-attention feature reflecting modality alignment enforced during generation. To enable faithful cross-generator evaluation, we further introduce MMDF, a new multimodal deepfake dataset spanning diverse manipulation types and rapidly evolving synthesis paradigms, including GANs, diffusion, and flow-matching. Extensive experiments demonstrate that X-AVDT achieves leading performance on MMDF and generalizes strongly to external benchmarks and unseen generators, outperforming existing methods with accuracy improved by 13.1%. Our findings highlight the importance of leveraging internal audio-visual consistency cues for robustness to future generators in deepfake detection.
title X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.08483