ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908899545710592 |
|---|---|
| author | Pan, Jianan Zhang, Yuanming Huang, Kejie |
| author_facet | Pan, Jianan Zhang, Yuanming Huang, Kejie |
| contents | Current keyword spotting systems primarily use phoneme-level matching to distinguish confusable words but ignore user-specific pronunciation traits like prosody (intonation, stress, rhythm). This paper presents ProKWS, a novel framework integrating fine-grained phoneme learning with personalized prosody modeling. We design a dual-stream encoder where one stream derives robust phonemic representations through contrastive learning, while the other extracts speaker-specific prosodic patterns. A collaborative fusion module dynamically combines phonemic and prosodic information, enhancing adaptability across acoustic environments. Experiments show ProKWS delivers highly competitive performance, comparable to state-of-the-art models on standard benchmarks and demonstrates strong robustness for personalized keywords with tone and intent variations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_18024 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody Pan, Jianan Zhang, Yuanming Huang, Kejie Audio and Speech Processing Artificial Intelligence Computation and Language Sound Current keyword spotting systems primarily use phoneme-level matching to distinguish confusable words but ignore user-specific pronunciation traits like prosody (intonation, stress, rhythm). This paper presents ProKWS, a novel framework integrating fine-grained phoneme learning with personalized prosody modeling. We design a dual-stream encoder where one stream derives robust phonemic representations through contrastive learning, while the other extracts speaker-specific prosodic patterns. A collaborative fusion module dynamically combines phonemic and prosodic information, enhancing adaptability across acoustic environments. Experiments show ProKWS delivers highly competitive performance, comparable to state-of-the-art models on standard benchmarks and demonstrates strong robustness for personalized keywords with tone and intent variations. |
| title | ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody |
| topic | Audio and Speech Processing Artificial Intelligence Computation and Language Sound |
| url | https://arxiv.org/abs/2603.18024 |