GE2E-KWS: Generalized End-to-End Training and Evaluation for Zero-shot Keyword Spotting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Pai, Bartel, Jacob W., Agarwal, Dhruuv, Partridge, Kurt, Park, Hyun Jin, Wang, Quan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918323229294592
author Zhu, Pai
Bartel, Jacob W.
Agarwal, Dhruuv
Partridge, Kurt
Park, Hyun Jin
Wang, Quan
author_facet Zhu, Pai
Bartel, Jacob W.
Agarwal, Dhruuv
Partridge, Kurt
Park, Hyun Jin
Wang, Quan
contents We propose GE2E-KWS -- a generalized end-to-end training and evaluation framework for customized keyword spotting. Specifically, enrollment utterances are separated and grouped by keywords from the training batch and their embedding centroids are compared to all other test utterance embeddings to compute the loss. This simulates runtime enrollment and verification stages, and improves convergence stability and training speed by optimizing matrix operations compared to SOTA triplet loss approaches. To benchmark different models reliably, we propose an evaluation process that mimics the production environment and compute metrics that directly measure keyword matching accuracy. Trained with GE2E loss, our 419KB quantized conformer model beats a 7.5GB ASR encoder by 23.6% relative AUC, and beats a same size triplet loss model by 60.7% AUC. Our KWS models are natively streamable with low memory footprints, and designed to continuously run on-device with no retraining needed for new keywords (zero-shot).
format Preprint
id arxiv_https___arxiv_org_abs_2410_16647
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GE2E-KWS: Generalized End-to-End Training and Evaluation for Zero-shot Keyword Spotting
Zhu, Pai
Bartel, Jacob W.
Agarwal, Dhruuv
Partridge, Kurt
Park, Hyun Jin
Wang, Quan
Audio and Speech Processing
Artificial Intelligence
Machine Learning
We propose GE2E-KWS -- a generalized end-to-end training and evaluation framework for customized keyword spotting. Specifically, enrollment utterances are separated and grouped by keywords from the training batch and their embedding centroids are compared to all other test utterance embeddings to compute the loss. This simulates runtime enrollment and verification stages, and improves convergence stability and training speed by optimizing matrix operations compared to SOTA triplet loss approaches. To benchmark different models reliably, we propose an evaluation process that mimics the production environment and compute metrics that directly measure keyword matching accuracy. Trained with GE2E loss, our 419KB quantized conformer model beats a 7.5GB ASR encoder by 23.6% relative AUC, and beats a same size triplet loss model by 60.7% AUC. Our KWS models are natively streamable with low memory footprints, and designed to continuously run on-device with no retraining needed for new keywords (zero-shot).
title GE2E-KWS: Generalized End-to-End Training and Evaluation for Zero-shot Keyword Spotting
topic Audio and Speech Processing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.16647