STOPA: A Database of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Firc, Anton, Chhibber, Manasi, Mishra, Jagabandhu, Singh, Vishwanath Pratap, Kinnunen, Tomi, Malinka, Kamil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916997650972672
author Firc, Anton
Chhibber, Manasi
Mishra, Jagabandhu
Singh, Vishwanath Pratap
Kinnunen, Tomi
Malinka, Kamil
author_facet Firc, Anton
Chhibber, Manasi
Mishra, Jagabandhu
Singh, Vishwanath Pratap
Kinnunen, Tomi
Malinka, Kamil
contents A key research area in deepfake speech detection is source tracing - determining the origin of synthesised utterances. The approaches may involve identifying the acoustic model (AM), vocoder model (VM), or other generation-specific parameters. However, progress is limited by the lack of a dedicated, systematically curated dataset. To address this, we introduce STOPA, a systematically varied and metadata-rich dataset for deepfake speech source tracing, covering 8 AMs, 6 VMs, and diverse parameter settings across 700k samples from 13 distinct synthesisers. Unlike existing datasets, which often feature limited variation or sparse metadata, STOPA provides a systematically controlled framework covering a broader range of generative factors, such as the choice of the vocoder model, acoustic model, or pretrained weights, ensuring higher attribution reliability. This control improves attribution accuracy, aiding forensic analysis, deepfake detection, and generative model transparency.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19644
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STOPA: A Database of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution
Firc, Anton
Chhibber, Manasi
Mishra, Jagabandhu
Singh, Vishwanath Pratap
Kinnunen, Tomi
Malinka, Kamil
Sound
Artificial Intelligence
Cryptography and Security
Audio and Speech Processing
68T45, 68T10, 94A08
I.2.7; I.5.4; K.4.1
A key research area in deepfake speech detection is source tracing - determining the origin of synthesised utterances. The approaches may involve identifying the acoustic model (AM), vocoder model (VM), or other generation-specific parameters. However, progress is limited by the lack of a dedicated, systematically curated dataset. To address this, we introduce STOPA, a systematically varied and metadata-rich dataset for deepfake speech source tracing, covering 8 AMs, 6 VMs, and diverse parameter settings across 700k samples from 13 distinct synthesisers. Unlike existing datasets, which often feature limited variation or sparse metadata, STOPA provides a systematically controlled framework covering a broader range of generative factors, such as the choice of the vocoder model, acoustic model, or pretrained weights, ensuring higher attribution reliability. This control improves attribution accuracy, aiding forensic analysis, deepfake detection, and generative model transparency.
title STOPA: A Database of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution
topic Sound
Artificial Intelligence
Cryptography and Security
Audio and Speech Processing
68T45, 68T10, 94A08
I.2.7; I.5.4; K.4.1
url https://arxiv.org/abs/2505.19644