Dynamic Recognition of Speakers for Consent Management by Contrastive Embedding Replay

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shahmansoori, Arash, Roedig, Utz
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913564102492160
author Shahmansoori, Arash
Roedig, Utz
author_facet Shahmansoori, Arash
Roedig, Utz
contents Voice assistants overhear conversations and a consent management mechanism is required. Consent management can be implemented using speaker recognition. Users that do not give consent enrol their voice and all their further recordings are discarded. Building speaker recognition-based consent management is challenging as dynamic registration, removal, and re-registration of speakers must be efficiently handled. This work proposes a consent management system addressing the aforementioned challenges. A contrastive based training is applied to learn the underlying speaker equivariance inductive bias. The contrastive features for buckets of speakers are trained a few steps into each iteration and act as replay buffers. These features are progressively selected using a multi-strided random sampler for classification. Moreover, new methods for dynamic registration using a portion of old utterances, removal, and re-registration of speakers are proposed. The results verify memory efficiency and dynamic capabilities of the proposed methods and outperform the existing approach from the literature.
format Preprint
id arxiv_https___arxiv_org_abs_2205_08459
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Dynamic Recognition of Speakers for Consent Management by Contrastive Embedding Replay
Shahmansoori, Arash
Roedig, Utz
Sound
Machine Learning
Audio and Speech Processing
Voice assistants overhear conversations and a consent management mechanism is required. Consent management can be implemented using speaker recognition. Users that do not give consent enrol their voice and all their further recordings are discarded. Building speaker recognition-based consent management is challenging as dynamic registration, removal, and re-registration of speakers must be efficiently handled. This work proposes a consent management system addressing the aforementioned challenges. A contrastive based training is applied to learn the underlying speaker equivariance inductive bias. The contrastive features for buckets of speakers are trained a few steps into each iteration and act as replay buffers. These features are progressively selected using a multi-strided random sampler for classification. Moreover, new methods for dynamic registration using a portion of old utterances, removal, and re-registration of speakers are proposed. The results verify memory efficiency and dynamic capabilities of the proposed methods and outperform the existing approach from the literature.
title Dynamic Recognition of Speakers for Consent Management by Contrastive Embedding Replay
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2205.08459