SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shahzad, Sahibzada Adil, Hashmi, Ammarah, Yamagishi, Junichi, Yasuda, Yusuke, Tsao, Yu, Lin, Chia-Wen, Peng, Yan-Tsung, Wang, Hsin-Min
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915892352253952
author Shahzad, Sahibzada Adil
Hashmi, Ammarah
Yamagishi, Junichi
Yasuda, Yusuke
Tsao, Yu
Lin, Chia-Wen
Peng, Yan-Tsung
Wang, Hsin-Min
author_facet Shahzad, Sahibzada Adil
Hashmi, Ammarah
Yamagishi, Junichi
Yasuda, Yusuke
Tsao, Yu
Lin, Chia-Wen
Peng, Yan-Tsung
Wang, Hsin-Min
contents Multimodal deepfakes can exhibit subtle visual artifacts and cross-modal inconsistencies, which remain challenging to detect, especially when detectors are trained primarily on curated synthetic forgeries. Such synthetic dependence can introduce dataset and generator bias, limiting scalability and robustness to unseen manipulations. We propose SAVe, a self-supervised audio-visual deepfake detection framework that learns entirely on authentic videos. SAVe generates on-the-fly, identity-preserving, region-aware self-blended pseudo-manipulations to emulate tampering artifacts, enabling the model to learn complementary visual cues across multiple facial granularities. To capture cross-modal evidence, SAVe also models lip-speech synchronization via an audio-visual alignment component that detects temporal misalignment patterns characteristic of audio-visual forgeries. Experiments on FakeAVCeleb and AV-LipSync-TIMIT demonstrate competitive in-domain performance and strong cross-dataset generalization, highlighting self-supervised learning as a scalable paradigm for multimodal deepfake detection.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25140
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
Shahzad, Sahibzada Adil
Hashmi, Ammarah
Yamagishi, Junichi
Yasuda, Yusuke
Tsao, Yu
Lin, Chia-Wen
Peng, Yan-Tsung
Wang, Hsin-Min
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
Sound
Multimodal deepfakes can exhibit subtle visual artifacts and cross-modal inconsistencies, which remain challenging to detect, especially when detectors are trained primarily on curated synthetic forgeries. Such synthetic dependence can introduce dataset and generator bias, limiting scalability and robustness to unseen manipulations. We propose SAVe, a self-supervised audio-visual deepfake detection framework that learns entirely on authentic videos. SAVe generates on-the-fly, identity-preserving, region-aware self-blended pseudo-manipulations to emulate tampering artifacts, enabling the model to learn complementary visual cues across multiple facial granularities. To capture cross-modal evidence, SAVe also models lip-speech synchronization via an audio-visual alignment component that detects temporal misalignment patterns characteristic of audio-visual forgeries. Experiments on FakeAVCeleb and AV-LipSync-TIMIT demonstrate competitive in-domain performance and strong cross-dataset generalization, highlighting self-supervised learning as a scalable paradigm for multimodal deepfake detection.
title SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
Sound
url https://arxiv.org/abs/2603.25140