ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Jianan, Zhang, Yuanming, Huang, Kejie
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908899545710592
author Pan, Jianan
Zhang, Yuanming
Huang, Kejie
author_facet Pan, Jianan
Zhang, Yuanming
Huang, Kejie
contents Current keyword spotting systems primarily use phoneme-level matching to distinguish confusable words but ignore user-specific pronunciation traits like prosody (intonation, stress, rhythm). This paper presents ProKWS, a novel framework integrating fine-grained phoneme learning with personalized prosody modeling. We design a dual-stream encoder where one stream derives robust phonemic representations through contrastive learning, while the other extracts speaker-specific prosodic patterns. A collaborative fusion module dynamically combines phonemic and prosodic information, enhancing adaptability across acoustic environments. Experiments show ProKWS delivers highly competitive performance, comparable to state-of-the-art models on standard benchmarks and demonstrates strong robustness for personalized keywords with tone and intent variations.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18024
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody
Pan, Jianan
Zhang, Yuanming
Huang, Kejie
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
Current keyword spotting systems primarily use phoneme-level matching to distinguish confusable words but ignore user-specific pronunciation traits like prosody (intonation, stress, rhythm). This paper presents ProKWS, a novel framework integrating fine-grained phoneme learning with personalized prosody modeling. We design a dual-stream encoder where one stream derives robust phonemic representations through contrastive learning, while the other extracts speaker-specific prosodic patterns. A collaborative fusion module dynamically combines phonemic and prosodic information, enhancing adaptability across acoustic environments. Experiments show ProKWS delivers highly competitive performance, comparable to state-of-the-art models on standard benchmarks and demonstrates strong robustness for personalized keywords with tone and intent variations.
title ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
url https://arxiv.org/abs/2603.18024