Explaining Speaker and Spoof Embeddings via Probing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Xuechen, Yamagishi, Junichi, Sahidullah, Md, kinnunen, Tomi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909440219807744
author Liu, Xuechen
Yamagishi, Junichi
Sahidullah, Md
kinnunen, Tomi
author_facet Liu, Xuechen
Yamagishi, Junichi
Sahidullah, Md
kinnunen, Tomi
contents This study investigates the explainability of embedding representations, specifically those used in modern audio spoofing detection systems based on deep neural networks, known as spoof embeddings. Building on established work in speaker embedding explainability, we examine how well these spoof embeddings capture speaker-related information. We train simple neural classifiers using either speaker or spoof embeddings as input, with speaker-related attributes as target labels. These attributes are categorized into two groups: metadata-based traits (e.g., gender, age) and acoustic traits (e.g., fundamental frequency, speaking rate). Our experiments on the ASVspoof 2019 LA evaluation set demonstrate that spoof embeddings preserve several key traits, including gender, speaking rate, F0, and duration. Further analysis of gender and speaking rate indicates that the spoofing detector partially preserves these traits, potentially to ensure the decision process remains robust against them.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18191
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Explaining Speaker and Spoof Embeddings via Probing
Liu, Xuechen
Yamagishi, Junichi
Sahidullah, Md
kinnunen, Tomi
Sound
Audio and Speech Processing
This study investigates the explainability of embedding representations, specifically those used in modern audio spoofing detection systems based on deep neural networks, known as spoof embeddings. Building on established work in speaker embedding explainability, we examine how well these spoof embeddings capture speaker-related information. We train simple neural classifiers using either speaker or spoof embeddings as input, with speaker-related attributes as target labels. These attributes are categorized into two groups: metadata-based traits (e.g., gender, age) and acoustic traits (e.g., fundamental frequency, speaking rate). Our experiments on the ASVspoof 2019 LA evaluation set demonstrate that spoof embeddings preserve several key traits, including gender, speaking rate, F0, and duration. Further analysis of gender and speaking rate indicates that the spoofing detector partially preserves these traits, potentially to ensure the decision process remains robust against them.
title Explaining Speaker and Spoof Embeddings via Probing
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.18191