Source Tracing of Synthetic Speech Systems Through Paralinguistic Pre-Trained Representations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Girish, Akhtar, Mohd Mujtaba, Phukan, Orchid Chetia, Singh, Drishti, Behera, Swarup Ranjan, Reddy, Pailla Balakrishna, Buduru, Arun Balaji, Sharma, Rajesh
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910979989700608
author Girish
Akhtar, Mohd Mujtaba
Phukan, Orchid Chetia
Singh, Drishti
Behera, Swarup Ranjan
Reddy, Pailla Balakrishna
Buduru, Arun Balaji
Sharma, Rajesh
author_facet Girish
Akhtar, Mohd Mujtaba
Phukan, Orchid Chetia
Singh, Drishti
Behera, Swarup Ranjan
Reddy, Pailla Balakrishna
Buduru, Arun Balaji
Sharma, Rajesh
contents In this work, we focus on source tracing of synthetic speech generation systems (STSGS). Each source embeds distinctive paralinguistic features--such as pitch, tone, rhythm, and intonation--into their synthesized speech, reflecting the underlying design of the generation model. While previous research has explored representations from speech pre-trained models (SPTMs), the use of representations from SPTM pre-trained for paralinguistic speech processing, which excel in paralinguistic tasks like synthetic speech detection, speech emotion recognition has not been investigated for STSGS. We hypothesize that representations from paralinguistic SPTM will be more effective due to its ability to capture source-specific paralinguistic cues attributing to its paralinguistic pre-training. Our comparative study of representations from various SOTA SPTMs, including paralinguistic, monolingual, multilingual, and speaker recognition, validates this hypothesis. Furthermore, we explore fusion of representations and propose TRIO, a novel framework that fuses SPTMs using a gated mechanism for adaptive weighting, followed by canonical correlation loss for inter-representation alignment and self-attention for feature refinement. By fusing TRILLsson (Paralinguistic SPTM) and x-vector (Speaker recognition SPTM), TRIO outperforms individual SPTMs, baseline fusion methods, and sets new SOTA for STSGS in comparison to previous works.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01157
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Source Tracing of Synthetic Speech Systems Through Paralinguistic Pre-Trained Representations
Girish
Akhtar, Mohd Mujtaba
Phukan, Orchid Chetia
Singh, Drishti
Behera, Swarup Ranjan
Reddy, Pailla Balakrishna
Buduru, Arun Balaji
Sharma, Rajesh
Audio and Speech Processing
Sound
In this work, we focus on source tracing of synthetic speech generation systems (STSGS). Each source embeds distinctive paralinguistic features--such as pitch, tone, rhythm, and intonation--into their synthesized speech, reflecting the underlying design of the generation model. While previous research has explored representations from speech pre-trained models (SPTMs), the use of representations from SPTM pre-trained for paralinguistic speech processing, which excel in paralinguistic tasks like synthetic speech detection, speech emotion recognition has not been investigated for STSGS. We hypothesize that representations from paralinguistic SPTM will be more effective due to its ability to capture source-specific paralinguistic cues attributing to its paralinguistic pre-training. Our comparative study of representations from various SOTA SPTMs, including paralinguistic, monolingual, multilingual, and speaker recognition, validates this hypothesis. Furthermore, we explore fusion of representations and propose TRIO, a novel framework that fuses SPTMs using a gated mechanism for adaptive weighting, followed by canonical correlation loss for inter-representation alignment and self-attention for feature refinement. By fusing TRILLsson (Paralinguistic SPTM) and x-vector (Speaker recognition SPTM), TRIO outperforms individual SPTMs, baseline fusion methods, and sets new SOTA for STSGS in comparison to previous works.
title Source Tracing of Synthetic Speech Systems Through Paralinguistic Pre-Trained Representations
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2506.01157