Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.03096 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909024722616320 |
|---|---|
| author | van Rensburg, Kyle Janse van Niekerk, Benjamin Kamper, Herman |
| author_facet | van Rensburg, Kyle Janse van Niekerk, Benjamin Kamper, Herman |
| contents | How do speech models trained through self-supervised learning structure their representations? Previous studies have looked at how information is encoded in feature vectors across different layers. But few studies have considered whether speech characteristics are captured within individual dimensions of SSL features. In this paper we specifically look at speaker information using PCA on utterance-averaged representations. For a range of SSL models, we find that the principal dimension that explains most variance encodes pitch and associated characteristics like gender. Other individual principal dimensions correlate with intensity, noise levels, the second formant, and higher frequency characteristics. We then use synthesis analyses to show that the dimensions for most characteristics are isolated from each other's influence. We further show that characteristics can be changed by manipulating the corresponding dimensions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_03096 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features van Rensburg, Kyle Janse van Niekerk, Benjamin Kamper, Herman Audio and Speech Processing Computation and Language How do speech models trained through self-supervised learning structure their representations? Previous studies have looked at how information is encoded in feature vectors across different layers. But few studies have considered whether speech characteristics are captured within individual dimensions of SSL features. In this paper we specifically look at speaker information using PCA on utterance-averaged representations. For a range of SSL models, we find that the principal dimension that explains most variance encodes pitch and associated characteristics like gender. Other individual principal dimensions correlate with intensity, noise levels, the second formant, and higher frequency characteristics. We then use synthesis analyses to show that the dimensions for most characteristics are isolated from each other's influence. We further show that characteristics can be changed by manipulating the corresponding dimensions. |
| title | Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features |
| topic | Audio and Speech Processing Computation and Language |
| url | https://arxiv.org/abs/2603.03096 |