Layer-aware TDNN: Speaker Recognition Using Multi-Layer Features from Pre-Trained Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kim, Jin Sob, Park, Hyun Joon, Shin, Wooseok, Yun, Juan, Han, Sung Won
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917145439371264
author Kim, Jin Sob
Park, Hyun Joon
Shin, Wooseok
Yun, Juan
Han, Sung Won
author_facet Kim, Jin Sob
Park, Hyun Joon
Shin, Wooseok
Yun, Juan
Han, Sung Won
contents Recent advances in self-supervised learning (SSL) on Transformers have significantly improved speaker verification (SV) by providing domain-general speech representations. However, existing approaches have underutilized the multi-layered nature of SSL encoders. To address this limitation, we propose the layer-aware time-delay neural network (L-TDNN), which directly performs layer/frame-wise processing on the layer-wise hidden state outputs from pre-trained models, extracting fixed-size speaker vectors. L-TDNN comprises a layer-aware convolutional network, a frame-adaptive layer aggregation, and attentive statistic pooling, explicitly modeling of the recognition and processing of previously overlooked layer dimension. We evaluated L-TDNN across multiple speech SSL Transformers and diverse speech-speaker corpora against other approaches for leveraging pre-trained encoders. L-TDNN consistently demonstrated robust verification performance, achieving the lowest error rates throughout the experiments. Concurrently, it stood out in terms of model compactness and exhibited inference efficiency comparable to the existing systems. These results highlight the advantages derived from the proposed layer-aware processing approach. Future work includes exploring joint training with SSL frontends and the incorporation of score calibration to further enhance state-of-the-art verification performance.
format Preprint
id arxiv_https___arxiv_org_abs_2409_07770
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Layer-aware TDNN: Speaker Recognition Using Multi-Layer Features from Pre-Trained Models
Kim, Jin Sob
Park, Hyun Joon
Shin, Wooseok
Yun, Juan
Han, Sung Won
Audio and Speech Processing
Artificial Intelligence
Recent advances in self-supervised learning (SSL) on Transformers have significantly improved speaker verification (SV) by providing domain-general speech representations. However, existing approaches have underutilized the multi-layered nature of SSL encoders. To address this limitation, we propose the layer-aware time-delay neural network (L-TDNN), which directly performs layer/frame-wise processing on the layer-wise hidden state outputs from pre-trained models, extracting fixed-size speaker vectors. L-TDNN comprises a layer-aware convolutional network, a frame-adaptive layer aggregation, and attentive statistic pooling, explicitly modeling of the recognition and processing of previously overlooked layer dimension. We evaluated L-TDNN across multiple speech SSL Transformers and diverse speech-speaker corpora against other approaches for leveraging pre-trained encoders. L-TDNN consistently demonstrated robust verification performance, achieving the lowest error rates throughout the experiments. Concurrently, it stood out in terms of model compactness and exhibited inference efficiency comparable to the existing systems. These results highlight the advantages derived from the proposed layer-aware processing approach. Future work includes exploring joint training with SSL frontends and the incorporation of score calibration to further enhance state-of-the-art verification performance.
title Layer-aware TDNN: Speaker Recognition Using Multi-Layer Features from Pre-Trained Models
topic Audio and Speech Processing
Artificial Intelligence
url https://arxiv.org/abs/2409.07770