GE2E-KWS: Generalized End-to-End Training and Evaluation for Zero-shot Keyword Spotting
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918323229294592 |
|---|---|
| author | Zhu, Pai Bartel, Jacob W. Agarwal, Dhruuv Partridge, Kurt Park, Hyun Jin Wang, Quan |
| author_facet | Zhu, Pai Bartel, Jacob W. Agarwal, Dhruuv Partridge, Kurt Park, Hyun Jin Wang, Quan |
| contents | We propose GE2E-KWS -- a generalized end-to-end training and evaluation framework for customized keyword spotting. Specifically, enrollment utterances are separated and grouped by keywords from the training batch and their embedding centroids are compared to all other test utterance embeddings to compute the loss. This simulates runtime enrollment and verification stages, and improves convergence stability and training speed by optimizing matrix operations compared to SOTA triplet loss approaches. To benchmark different models reliably, we propose an evaluation process that mimics the production environment and compute metrics that directly measure keyword matching accuracy. Trained with GE2E loss, our 419KB quantized conformer model beats a 7.5GB ASR encoder by 23.6% relative AUC, and beats a same size triplet loss model by 60.7% AUC. Our KWS models are natively streamable with low memory footprints, and designed to continuously run on-device with no retraining needed for new keywords (zero-shot). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_16647 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | GE2E-KWS: Generalized End-to-End Training and Evaluation for Zero-shot Keyword Spotting Zhu, Pai Bartel, Jacob W. Agarwal, Dhruuv Partridge, Kurt Park, Hyun Jin Wang, Quan Audio and Speech Processing Artificial Intelligence Machine Learning We propose GE2E-KWS -- a generalized end-to-end training and evaluation framework for customized keyword spotting. Specifically, enrollment utterances are separated and grouped by keywords from the training batch and their embedding centroids are compared to all other test utterance embeddings to compute the loss. This simulates runtime enrollment and verification stages, and improves convergence stability and training speed by optimizing matrix operations compared to SOTA triplet loss approaches. To benchmark different models reliably, we propose an evaluation process that mimics the production environment and compute metrics that directly measure keyword matching accuracy. Trained with GE2E loss, our 419KB quantized conformer model beats a 7.5GB ASR encoder by 23.6% relative AUC, and beats a same size triplet loss model by 60.7% AUC. Our KWS models are natively streamable with low memory footprints, and designed to continuously run on-device with no retraining needed for new keywords (zero-shot). |
| title | GE2E-KWS: Generalized End-to-End Training and Evaluation for Zero-shot Keyword Spotting |
| topic | Audio and Speech Processing Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2410.16647 |