Multi-Sample Dynamic Time Warping for Few-Shot Keyword Spotting

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wilkinghoff, Kevin, Cornaggia-Urrigshardt, Alessia
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916274230001664
author Wilkinghoff, Kevin
Cornaggia-Urrigshardt, Alessia
author_facet Wilkinghoff, Kevin
Cornaggia-Urrigshardt, Alessia
contents In multi-sample keyword spotting, each keyword class is represented by multiple spoken instances, called samples. A naïve approach to detect keywords in a target sequence consists of querying all samples of all classes using sub-sequence dynamic time warping. However, the resulting processing time increases linearly with respect to the number of samples belonging to each class. Alternatively, only a single Fréchet mean can be queried for each class, resulting in reduced processing time but usually also in worse detection performance as the variability of the query samples is not captured sufficiently well. In this work, multi-sample dynamic time warping is proposed to compute class-specific cost-tensors that include the variability of all query samples. To significantly reduce the computational complexity during inference, these cost tensors are converted to cost matrices before applying dynamic time warping. In experimental evaluations for few-shot keyword spotting, it is shown that this method yields a very similar performance as using all individual query samples as templates while having a runtime that is only slightly slower than when using Fréchet means.
format Preprint
id arxiv_https___arxiv_org_abs_2404_14903
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Sample Dynamic Time Warping for Few-Shot Keyword Spotting
Wilkinghoff, Kevin
Cornaggia-Urrigshardt, Alessia
Audio and Speech Processing
Information Retrieval
Sound
In multi-sample keyword spotting, each keyword class is represented by multiple spoken instances, called samples. A naïve approach to detect keywords in a target sequence consists of querying all samples of all classes using sub-sequence dynamic time warping. However, the resulting processing time increases linearly with respect to the number of samples belonging to each class. Alternatively, only a single Fréchet mean can be queried for each class, resulting in reduced processing time but usually also in worse detection performance as the variability of the query samples is not captured sufficiently well. In this work, multi-sample dynamic time warping is proposed to compute class-specific cost-tensors that include the variability of all query samples. To significantly reduce the computational complexity during inference, these cost tensors are converted to cost matrices before applying dynamic time warping. In experimental evaluations for few-shot keyword spotting, it is shown that this method yields a very similar performance as using all individual query samples as templates while having a runtime that is only slightly slower than when using Fréchet means.
title Multi-Sample Dynamic Time Warping for Few-Shot Keyword Spotting
topic Audio and Speech Processing
Information Retrieval
Sound
url https://arxiv.org/abs/2404.14903