Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sinha, Abhijit, Kumar, Harishankar, Joshi, Mohit, Kathania, Hemant Kumar, Narayanan, Shrikanth, Kadiri, Sudarsana Reddy
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912537396641792
author Sinha, Abhijit
Kumar, Harishankar
Joshi, Mohit
Kathania, Hemant Kumar
Narayanan, Shrikanth
Kadiri, Sudarsana Reddy
author_facet Sinha, Abhijit
Kumar, Harishankar
Joshi, Mohit
Kathania, Hemant Kumar
Narayanan, Shrikanth
Kadiri, Sudarsana Reddy
contents Children's speech presents challenges for age and gender classification due to high variability in pitch, articulation, and developmental traits. While self-supervised learning (SSL) models perform well on adult speech tasks, their ability to encode speaker traits in children remains underexplored. This paper presents a detailed layer-wise analysis of four Wav2Vec2 variants using the PFSTAR and CMU Kids datasets. Results show that early layers (1-7) capture speaker-specific cues more effectively than deeper layers, which increasingly focus on linguistic information. Applying PCA further improves classification, reducing redundancy and highlighting the most informative components. The Wav2Vec2-large-lv60 model achieves 97.14% (age) and 98.20% (gender) on CMU Kids; base-100h and large-lv60 models reach 86.05% and 95.00% on PFSTAR. These results reveal how speaker traits are structured across SSL model depth and support more targeted, adaptive strategies for child-aware speech interfaces.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10332
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
Sinha, Abhijit
Kumar, Harishankar
Joshi, Mohit
Kathania, Hemant Kumar
Narayanan, Shrikanth
Kadiri, Sudarsana Reddy
Audio and Speech Processing
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Sound
Children's speech presents challenges for age and gender classification due to high variability in pitch, articulation, and developmental traits. While self-supervised learning (SSL) models perform well on adult speech tasks, their ability to encode speaker traits in children remains underexplored. This paper presents a detailed layer-wise analysis of four Wav2Vec2 variants using the PFSTAR and CMU Kids datasets. Results show that early layers (1-7) capture speaker-specific cues more effectively than deeper layers, which increasingly focus on linguistic information. Applying PCA further improves classification, reducing redundancy and highlighting the most informative components. The Wav2Vec2-large-lv60 model achieves 97.14% (age) and 98.20% (gender) on CMU Kids; base-100h and large-lv60 models reach 86.05% and 95.00% on PFSTAR. These results reveal how speaker traits are structured across SSL model depth and support more targeted, adaptive strategies for child-aware speech interfaces.
title Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
topic Audio and Speech Processing
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Sound
url https://arxiv.org/abs/2508.10332