AuViRe: Audio-visual Speech Representation Reconstruction for Deepfake Temporal Localization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koutlis, Christos, Papadopoulos, Symeon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917100883279872
author Koutlis, Christos
Papadopoulos, Symeon
author_facet Koutlis, Christos
Papadopoulos, Symeon
contents With the rapid advancement of sophisticated synthetic audio-visual content, e.g., for subtle malicious manipulations, ensuring the integrity of digital media has become paramount. This work presents a novel approach to temporal localization of deepfakes by leveraging Audio-Visual Speech Representation Reconstruction (AuViRe). Specifically, our approach reconstructs speech representations from one modality (e.g., lip movements) based on the other (e.g., audio waveform). Cross-modal reconstruction is significantly more challenging in manipulated video segments, leading to amplified discrepancies, thereby providing robust discriminative cues for precise temporal forgery localization. AuViRe outperforms the state of the art by +8.9 AP@0.95 on LAV-DF, +9.6 AP@0.5 on AV-Deepfake1M, and +5.1 AUC on an in-the-wild experiment. Code available at https://github.com/mever-team/auvire.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18993
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AuViRe: Audio-visual Speech Representation Reconstruction for Deepfake Temporal Localization
Koutlis, Christos
Papadopoulos, Symeon
Computer Vision and Pattern Recognition
With the rapid advancement of sophisticated synthetic audio-visual content, e.g., for subtle malicious manipulations, ensuring the integrity of digital media has become paramount. This work presents a novel approach to temporal localization of deepfakes by leveraging Audio-Visual Speech Representation Reconstruction (AuViRe). Specifically, our approach reconstructs speech representations from one modality (e.g., lip movements) based on the other (e.g., audio waveform). Cross-modal reconstruction is significantly more challenging in manipulated video segments, leading to amplified discrepancies, thereby providing robust discriminative cues for precise temporal forgery localization. AuViRe outperforms the state of the art by +8.9 AP@0.95 on LAV-DF, +9.6 AP@0.5 on AV-Deepfake1M, and +5.1 AUC on an in-the-wild experiment. Code available at https://github.com/mever-team/auvire.
title AuViRe: Audio-visual Speech Representation Reconstruction for Deepfake Temporal Localization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.18993