LENS-DF: Deepfake Detection and Temporal Localization for Long-Form Noisy Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Xuechen, Ge, Wanying, Wang, Xin, Yamagishi, Junichi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908463561441280
author Liu, Xuechen
Ge, Wanying
Wang, Xin
Yamagishi, Junichi
author_facet Liu, Xuechen
Ge, Wanying
Wang, Xin
Yamagishi, Junichi
contents This study introduces LENS-DF, a novel and comprehensive recipe for training and evaluating audio deepfake detection and temporal localization under complicated and realistic audio conditions. The generation part of the recipe outputs audios from the input dataset with several critical characteristics, such as longer duration, noisy conditions, and containing multiple speakers, in a controllable fashion. The corresponding detection and localization protocol uses models. We conduct experiments based on self-supervised learning front-end and simple back-end. The results indicate that models trained using data generated with LENS-DF consistently outperform those trained via conventional recipes, demonstrating the effectiveness and usefulness of LENS-DF for robust audio deepfake detection and localization. We also conduct ablation studies on the variations introduced, investigating their impact on and relevance to realistic challenges in the field.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16220
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LENS-DF: Deepfake Detection and Temporal Localization for Long-Form Noisy Speech
Liu, Xuechen
Ge, Wanying
Wang, Xin
Yamagishi, Junichi
Sound
Cryptography and Security
Audio and Speech Processing
This study introduces LENS-DF, a novel and comprehensive recipe for training and evaluating audio deepfake detection and temporal localization under complicated and realistic audio conditions. The generation part of the recipe outputs audios from the input dataset with several critical characteristics, such as longer duration, noisy conditions, and containing multiple speakers, in a controllable fashion. The corresponding detection and localization protocol uses models. We conduct experiments based on self-supervised learning front-end and simple back-end. The results indicate that models trained using data generated with LENS-DF consistently outperform those trained via conventional recipes, demonstrating the effectiveness and usefulness of LENS-DF for robust audio deepfake detection and localization. We also conduct ablation studies on the variations introduced, investigating their impact on and relevance to realistic challenges in the field.
title LENS-DF: Deepfake Detection and Temporal Localization for Long-Form Noisy Speech
topic Sound
Cryptography and Security
Audio and Speech Processing
url https://arxiv.org/abs/2507.16220