AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912384346488832 |
|---|---|
| author | Xiao, Yang Peng, Tianyi Zhou, Yanghao Das, Rohan Kumar |
| author_facet | Xiao, Yang Peng, Tianyi Zhou, Yanghao Das, Rohan Kumar |
| contents | Spoken keyword spotting (KWS) aims to identify keywords in audio for wide applications, especially on edge devices. Current small-footprint KWS systems focus on efficient model designs. However, their inference performance can decline in unseen environments or noisy backgrounds. Test-time adaptation (TTA) helps models adapt to test samples without needing the original training data. In this study, we present AdaKWS, the first TTA method for robust KWS to the best of our knowledge. Specifically, 1) We initially optimize the model's confidence by selecting reliable samples based on prediction entropy minimization and adjusting the normalization statistics in each batch. 2) We introduce pseudo-keyword consistency (PKC) to identify critical, reliable features without overfitting to noise. Our experiments show that AdaKWS outperforms other methods across various conditions, including Gaussian noise and real-scenario noises. The code will be released in due course. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_14600 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation Xiao, Yang Peng, Tianyi Zhou, Yanghao Das, Rohan Kumar Audio and Speech Processing Sound Spoken keyword spotting (KWS) aims to identify keywords in audio for wide applications, especially on edge devices. Current small-footprint KWS systems focus on efficient model designs. However, their inference performance can decline in unseen environments or noisy backgrounds. Test-time adaptation (TTA) helps models adapt to test samples without needing the original training data. In this study, we present AdaKWS, the first TTA method for robust KWS to the best of our knowledge. Specifically, 1) We initially optimize the model's confidence by selecting reliable samples based on prediction entropy minimization and adjusting the normalization statistics in each batch. 2) We introduce pseudo-keyword consistency (PKC) to identify critical, reliable features without overfitting to noise. Our experiments show that AdaKWS outperforms other methods across various conditions, including Gaussian noise and real-scenario noises. The code will be released in due course. |
| title | AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2505.14600 |