Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bora, Maheswar, Dhamija, Tashvik, Reddy, Shukesh, Chopin, Baptiste, Balaji, Pranav, Das, Abhijit, Dantcheva, Antitza
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917110357164032
author Bora, Maheswar
Dhamija, Tashvik
Reddy, Shukesh
Chopin, Baptiste
Balaji, Pranav
Das, Abhijit
Dantcheva, Antitza
author_facet Bora, Maheswar
Dhamija, Tashvik
Reddy, Shukesh
Chopin, Baptiste
Balaji, Pranav
Das, Abhijit
Dantcheva, Antitza
contents Deepfake generation has witnessed remarkable progress, contributing to highly realistic generated images, videos, and audio. While technically intriguing, such progress has raised serious concerns related to the misuse of manipulated media. To mitigate such misuse, robust and reliable deepfake detection is urgently needed. Towards this, we propose a novel network FauxNet, which is based on pre-trained Visual Speech Recognition (VSR) features. By extracting temporal VSR features from videos, we identify and segregate real videos from manipulated ones. The holy grail in this context has to do with zero-shot detection, i.e., generalizable detection, which we focus on in this work. FauxNet consistently outperforms the state-of-the-art in this setting. In addition, FauxNet is able to attribute - distinguish between generation techniques from which the videos stem. Finally, we propose new datasets, referred to as Authentica-Vox and Authentica-HDTF, comprising about 38,000 real and fake videos in total, the latter created with six recent deepfake generation techniques. We provide extensive analysis and results on the Authentica datasets and FaceForensics++, demonstrating the superiority of FauxNet. The Authentica datasets will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22443
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
Bora, Maheswar
Dhamija, Tashvik
Reddy, Shukesh
Chopin, Baptiste
Balaji, Pranav
Das, Abhijit
Dantcheva, Antitza
Computer Vision and Pattern Recognition
Deepfake generation has witnessed remarkable progress, contributing to highly realistic generated images, videos, and audio. While technically intriguing, such progress has raised serious concerns related to the misuse of manipulated media. To mitigate such misuse, robust and reliable deepfake detection is urgently needed. Towards this, we propose a novel network FauxNet, which is based on pre-trained Visual Speech Recognition (VSR) features. By extracting temporal VSR features from videos, we identify and segregate real videos from manipulated ones. The holy grail in this context has to do with zero-shot detection, i.e., generalizable detection, which we focus on in this work. FauxNet consistently outperforms the state-of-the-art in this setting. In addition, FauxNet is able to attribute - distinguish between generation techniques from which the videos stem. Finally, we propose new datasets, referred to as Authentica-Vox and Authentica-HDTF, comprising about 38,000 real and fake videos in total, the latter created with six recent deepfake generation techniques. We provide extensive analysis and results on the Authentica datasets and FaceForensics++, demonstrating the superiority of FauxNet. The Authentica datasets will be made publicly available.
title Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.22443