Referee: Reference-aware Audiovisual Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Boo, Hyemin, Lee, Eunsang, Lee, Jiyoung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915858263048192
author Boo, Hyemin
Lee, Eunsang
Lee, Jiyoung
author_facet Boo, Hyemin
Lee, Eunsang
Lee, Jiyoung
contents Deepfakes generated by advanced generative models have rapidly posed serious threats, yet existing audiovisual deepfake detection approaches struggle to generalize to unseen manipulation methods. To address this, we propose a novel reference-aware audiovisual deepfake detection method, called Referee to capture fine-grained identity discrepancies. Unlike existing methods that overfit to transient spatiotemporal artifacts, Referee employs identity bottleneck and matching modules to model the relational consistency of speaker-specific cues captured by a single one-shot example as a biometric anchor. Extensive experiments on FakeAVCeleb, FaceForensics++, and KoDF demonstrate that Referee achieves state-of-the-art results on cross-dataset and cross-language evaluation protocols, including a 99.4% AUC on KoDF. These results highlight that explicitly correlating reference-based biometric priors is a key frontier for achieving generalized and reliable audiovisual forensics. The code is available at https://github.com/ewha-mmai/referee.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27475
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Referee: Reference-aware Audiovisual Deepfake Detection
Boo, Hyemin
Lee, Eunsang
Lee, Jiyoung
Computer Vision and Pattern Recognition
Multimedia
Deepfakes generated by advanced generative models have rapidly posed serious threats, yet existing audiovisual deepfake detection approaches struggle to generalize to unseen manipulation methods. To address this, we propose a novel reference-aware audiovisual deepfake detection method, called Referee to capture fine-grained identity discrepancies. Unlike existing methods that overfit to transient spatiotemporal artifacts, Referee employs identity bottleneck and matching modules to model the relational consistency of speaker-specific cues captured by a single one-shot example as a biometric anchor. Extensive experiments on FakeAVCeleb, FaceForensics++, and KoDF demonstrate that Referee achieves state-of-the-art results on cross-dataset and cross-language evaluation protocols, including a 99.4% AUC on KoDF. These results highlight that explicitly correlating reference-based biometric priors is a key frontier for achieving generalized and reliable audiovisual forensics. The code is available at https://github.com/ewha-mmai/referee.
title Referee: Reference-aware Audiovisual Deepfake Detection
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2510.27475