Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Miseul, Park, Soo Jin, Byun, Kyungguen, Shin, Hyeon-Kyeong, Moon, Sunkuk, Zhang, Shuhua, Visser, Erik
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909794940485632
author Kim, Miseul
Park, Soo Jin
Byun, Kyungguen
Shin, Hyeon-Kyeong
Moon, Sunkuk
Zhang, Shuhua
Visser, Erik
author_facet Kim, Miseul
Park, Soo Jin
Byun, Kyungguen
Shin, Hyeon-Kyeong
Moon, Sunkuk
Zhang, Shuhua
Visser, Erik
contents Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speaker to be misclassified as different individuals, for example, when one raises their voice or speaks faster during conversation. To address this, we propose a style-controllable speech generation model that augments speech across diverse styles while preserving the target speaker's identity. The proposed system starts with diarized segments from a conventional diarizer. For each diarized segment, it generates augmented speech samples enriched with phonetic and stylistic diversity. And then, speaker embeddings from both the original and generated audio are blended to enhance the system's robustness in grouping segments with high intrinsic intra-speaker variability. We validate our approach on a simulated emotional speech dataset and the truncated AMI dataset, demonstrating significant improvements, with error rate reductions of 49% and 35% on each dataset, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14632
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
Kim, Miseul
Park, Soo Jin
Byun, Kyungguen
Shin, Hyeon-Kyeong
Moon, Sunkuk
Zhang, Shuhua
Visser, Erik
Audio and Speech Processing
Artificial Intelligence
Signal Processing
Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speaker to be misclassified as different individuals, for example, when one raises their voice or speaks faster during conversation. To address this, we propose a style-controllable speech generation model that augments speech across diverse styles while preserving the target speaker's identity. The proposed system starts with diarized segments from a conventional diarizer. For each diarized segment, it generates augmented speech samples enriched with phonetic and stylistic diversity. And then, speaker embeddings from both the original and generated audio are blended to enhance the system's robustness in grouping segments with high intrinsic intra-speaker variability. We validate our approach on a simulated emotional speech dataset and the truncated AMI dataset, demonstrating significant improvements, with error rate reductions of 49% and 35% on each dataset, respectively.
title Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
topic Audio and Speech Processing
Artificial Intelligence
Signal Processing
url https://arxiv.org/abs/2509.14632