Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Klein, Nicholas, Tak, Hemlata, Fullwood, James, Regmi, Krishna, Spinoulas, Leonidas, Sivaraman, Ganesh, Chen, Tianxiang, Khoury, Elie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917041089282048
author Klein, Nicholas
Tak, Hemlata
Fullwood, James
Regmi, Krishna
Spinoulas, Leonidas
Sivaraman, Ganesh
Chen, Tianxiang
Khoury, Elie
author_facet Klein, Nicholas
Tak, Hemlata
Fullwood, James
Regmi, Krishna
Spinoulas, Leonidas
Sivaraman, Ganesh
Chen, Tianxiang
Khoury, Elie
contents The field of visual and audio generation is burgeoning with new state-of-the-art methods. This rapid proliferation of new techniques underscores the need for robust solutions for detecting synthetic content in videos. In particular, when fine-grained alterations via localized manipulations are performed in visual, audio, or both domains, these subtle modifications add challenges to the detection algorithms. This paper presents solutions for the problems of deepfake video classification and localization. The methods were submitted to the ACM 1M Deepfakes Detection Challenge, achieving the best performance in the temporal localization task and a top four ranking in the classification task for the TestA split of the evaluation dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08141
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization
Klein, Nicholas
Tak, Hemlata
Fullwood, James
Regmi, Krishna
Spinoulas, Leonidas
Sivaraman, Ganesh
Chen, Tianxiang
Khoury, Elie
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
The field of visual and audio generation is burgeoning with new state-of-the-art methods. This rapid proliferation of new techniques underscores the need for robust solutions for detecting synthetic content in videos. In particular, when fine-grained alterations via localized manipulations are performed in visual, audio, or both domains, these subtle modifications add challenges to the detection algorithms. This paper presents solutions for the problems of deepfake video classification and localization. The methods were submitted to the ACM 1M Deepfakes Detection Challenge, achieving the best performance in the temporal localization task and a top four ranking in the classification task for the TestA split of the evaluation dataset.
title Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization
topic Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2508.08141