Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911393276493824 |
|---|---|
| author | Serre, Thomas Fontaine, Mathieu Benhaim, Éric Essid, Slim |
| author_facet | Serre, Thomas Fontaine, Mathieu Benhaim, Éric Essid, Slim |
| contents | Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorporate a representation of the target voice within the enhancement system, which is extracted from an enrollment clip of the target voice with upstream models. Those models are generally heavy as the speaker embedding's quality directly affects PSE performances. Yet, embeddings generated beforehand cannot account for the variations of the target voice during inference time. In this paper, we propose to perform on-thefly refinement of the speaker embedding using a tiny speaker encoder. We first introduce a novel contrastive knowledge distillation methodology in order to train a 150k-parameter encoder from complex embeddings. We then use this encoder within the enhancement system during inference and show that the proposed method greatly improves PSE performances while maintaining a low computational load. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_16235 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement Serre, Thomas Fontaine, Mathieu Benhaim, Éric Essid, Slim Sound Audio and Speech Processing Signal Processing Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorporate a representation of the target voice within the enhancement system, which is extracted from an enrollment clip of the target voice with upstream models. Those models are generally heavy as the speaker embedding's quality directly affects PSE performances. Yet, embeddings generated beforehand cannot account for the variations of the target voice during inference time. In this paper, we propose to perform on-thefly refinement of the speaker embedding using a tiny speaker encoder. We first introduce a novel contrastive knowledge distillation methodology in order to train a 150k-parameter encoder from complex embeddings. We then use this encoder within the enhancement system during inference and show that the proposed method greatly improves PSE performances while maintaining a low computational load. |
| title | Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement |
| topic | Sound Audio and Speech Processing Signal Processing |
| url | https://arxiv.org/abs/2601.16235 |