GraphemeAug: A Systematic Approach to Synthesized Hard Negative Keyword Spotting Examples
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914306226913280 |
|---|---|
| author | Zhang, Harry Partridge, Kurt Zhu, Pai Chen, Neng Park, Hyun Jin Agarwal, Dhruuv Wang, Quan |
| author_facet | Zhang, Harry Partridge, Kurt Zhu, Pai Chen, Neng Park, Hyun Jin Agarwal, Dhruuv Wang, Quan |
| contents | Spoken Keyword Spotting (KWS) is the task of distinguishing between the presence and absence of a keyword in audio. The accuracy of a KWS model hinges on its ability to correctly classify examples close to the keyword and non-keyword boundary. These boundary examples are often scarce in training data, limiting model performance. In this paper, we propose a method to systematically generate adversarial examples close to the decision boundary by making insertion/deletion/substitution edits on the keyword's graphemes. We evaluate this technique on held-out data for a popular keyword and show that the technique improves AUC on a dataset of synthetic hard negatives by 61% while maintaining quality on positives and ambient negative audio data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_14814 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | GraphemeAug: A Systematic Approach to Synthesized Hard Negative Keyword Spotting Examples Zhang, Harry Partridge, Kurt Zhu, Pai Chen, Neng Park, Hyun Jin Agarwal, Dhruuv Wang, Quan Sound Computation and Language Audio and Speech Processing Spoken Keyword Spotting (KWS) is the task of distinguishing between the presence and absence of a keyword in audio. The accuracy of a KWS model hinges on its ability to correctly classify examples close to the keyword and non-keyword boundary. These boundary examples are often scarce in training data, limiting model performance. In this paper, we propose a method to systematically generate adversarial examples close to the decision boundary by making insertion/deletion/substitution edits on the keyword's graphemes. We evaluate this technique on held-out data for a popular keyword and show that the technique improves AUC on a dataset of synthetic hard negatives by 61% while maintaining quality on positives and ambient negative audio data. |
| title | GraphemeAug: A Systematic Approach to Synthesized Hard Negative Keyword Spotting Examples |
| topic | Sound Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2505.14814 |