Saved in:
Bibliographic Details
Main Authors: van Rensburg, Kyle Janse, van Niekerk, Benjamin, Kamper, Herman
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.03096
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909024722616320
author van Rensburg, Kyle Janse
van Niekerk, Benjamin
Kamper, Herman
author_facet van Rensburg, Kyle Janse
van Niekerk, Benjamin
Kamper, Herman
contents How do speech models trained through self-supervised learning structure their representations? Previous studies have looked at how information is encoded in feature vectors across different layers. But few studies have considered whether speech characteristics are captured within individual dimensions of SSL features. In this paper we specifically look at speaker information using PCA on utterance-averaged representations. For a range of SSL models, we find that the principal dimension that explains most variance encodes pitch and associated characteristics like gender. Other individual principal dimensions correlate with intensity, noise levels, the second formant, and higher frequency characteristics. We then use synthesis analyses to show that the dimensions for most characteristics are isolated from each other's influence. We further show that characteristics can be changed by manipulating the corresponding dimensions.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03096
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features
van Rensburg, Kyle Janse
van Niekerk, Benjamin
Kamper, Herman
Audio and Speech Processing
Computation and Language
How do speech models trained through self-supervised learning structure their representations? Previous studies have looked at how information is encoded in feature vectors across different layers. But few studies have considered whether speech characteristics are captured within individual dimensions of SSL features. In this paper we specifically look at speaker information using PCA on utterance-averaged representations. For a range of SSL models, we find that the principal dimension that explains most variance encodes pitch and associated characteristics like gender. Other individual principal dimensions correlate with intensity, noise levels, the second formant, and higher frequency characteristics. We then use synthesis analyses to show that the dimensions for most characteristics are isolated from each other's influence. We further show that characteristics can be changed by manipulating the corresponding dimensions.
title Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features
topic Audio and Speech Processing
Computation and Language
url https://arxiv.org/abs/2603.03096