AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Yang, Peng, Tianyi, Zhou, Yanghao, Das, Rohan Kumar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912384346488832
author Xiao, Yang
Peng, Tianyi
Zhou, Yanghao
Das, Rohan Kumar
author_facet Xiao, Yang
Peng, Tianyi
Zhou, Yanghao
Das, Rohan Kumar
contents Spoken keyword spotting (KWS) aims to identify keywords in audio for wide applications, especially on edge devices. Current small-footprint KWS systems focus on efficient model designs. However, their inference performance can decline in unseen environments or noisy backgrounds. Test-time adaptation (TTA) helps models adapt to test samples without needing the original training data. In this study, we present AdaKWS, the first TTA method for robust KWS to the best of our knowledge. Specifically, 1) We initially optimize the model's confidence by selecting reliable samples based on prediction entropy minimization and adjusting the normalization statistics in each batch. 2) We introduce pseudo-keyword consistency (PKC) to identify critical, reliable features without overfitting to noise. Our experiments show that AdaKWS outperforms other methods across various conditions, including Gaussian noise and real-scenario noises. The code will be released in due course.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14600
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation
Xiao, Yang
Peng, Tianyi
Zhou, Yanghao
Das, Rohan Kumar
Audio and Speech Processing
Sound
Spoken keyword spotting (KWS) aims to identify keywords in audio for wide applications, especially on edge devices. Current small-footprint KWS systems focus on efficient model designs. However, their inference performance can decline in unseen environments or noisy backgrounds. Test-time adaptation (TTA) helps models adapt to test samples without needing the original training data. In this study, we present AdaKWS, the first TTA method for robust KWS to the best of our knowledge. Specifically, 1) We initially optimize the model's confidence by selecting reliable samples based on prediction entropy minimization and adjusting the normalization statistics in each batch. 2) We introduce pseudo-keyword consistency (PKC) to identify critical, reliable features without overfitting to noise. Our experiments show that AdaKWS outperforms other methods across various conditions, including Gaussian noise and real-scenario noises. The code will be released in due course.
title AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2505.14600