How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Tianchi, Zhang, Lin, Das, Rohan Kumar, Ma, Yi, Tao, Ruijie, Li, Haizhou
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914823612137472
author Liu, Tianchi
Zhang, Lin
Das, Rohan Kumar
Ma, Yi
Tao, Ruijie
Li, Haizhou
author_facet Liu, Tianchi
Zhang, Lin
Das, Rohan Kumar
Ma, Yi
Tao, Ruijie
Li, Haizhou
contents Partially manipulating a sentence can greatly change its meaning. Recent work shows that countermeasures (CMs) trained on partially spoofed audio can effectively detect such spoofing. However, the current understanding of the decision-making process of CMs is limited. We utilize Grad-CAM and introduce a quantitative analysis metric to interpret CMs' decisions. We find that CMs prioritize the artifacts of transition regions created when concatenating bona fide and spoofed audio. This focus differs from that of CMs trained on fully spoofed audio, which concentrate on the pattern differences between bona fide and spoofed parts. Our further investigation explains the varying nature of CMs' focus while making correct or incorrect predictions. These insights provide a basis for the design of CM models and the creation of datasets. Moreover, this work lays a foundation of interpretability in the field of partial spoofed audio detection that has not been well explored previously.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02483
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?
Liu, Tianchi
Zhang, Lin
Das, Rohan Kumar
Ma, Yi
Tao, Ruijie
Li, Haizhou
Audio and Speech Processing
Artificial Intelligence
Sound
Partially manipulating a sentence can greatly change its meaning. Recent work shows that countermeasures (CMs) trained on partially spoofed audio can effectively detect such spoofing. However, the current understanding of the decision-making process of CMs is limited. We utilize Grad-CAM and introduce a quantitative analysis metric to interpret CMs' decisions. We find that CMs prioritize the artifacts of transition regions created when concatenating bona fide and spoofed audio. This focus differs from that of CMs trained on fully spoofed audio, which concentrate on the pattern differences between bona fide and spoofed parts. Our further investigation explains the varying nature of CMs' focus while making correct or incorrect predictions. These insights provide a basis for the design of CM models and the creation of datasets. Moreover, this work lays a foundation of interpretability in the field of partial spoofed audio detection that has not been well explored previously.
title How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2406.02483