Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Yuhang, Lin, Haopeng, Qian, Jiale, Yan, Ruiqi, Meng, Hao, Xie, Hanke, Wen, Hanlin, Yin, Shunshun, Tao, Ming, Chen, Xie, Xie, Lei, Wang, Xinsheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918412538609664
author Dai, Yuhang
Lin, Haopeng
Qian, Jiale
Yan, Ruiqi
Meng, Hao
Xie, Hanke
Wen, Hanlin
Yin, Shunshun
Tao, Ming
Chen, Xie
Xie, Lei
Wang, Xinsheng
author_facet Dai, Yuhang
Lin, Haopeng
Qian, Jiale
Yan, Ruiqi
Meng, Hao
Xie, Hanke
Wen, Hanlin
Yin, Shunshun
Tao, Ming
Chen, Xie
Xie, Lei
Wang, Xinsheng
contents Large Audio-Language Models (LALMs) have demonstrated remarkable performance in end-to-end speaker diarization and recognition. However, their speaker discriminability remains limited due to the scarcity of large-scale conversational data and the absence of explicit speaker representation optimization. To address this, we propose GLSC-SDR, a paradigm that jointly trains speaker classification with diarization and recognition. We further introduce a Global-Local Speaker Classification strategy, which uses clustered speakers as global labels and re-encoded intra-cluster speakers as local labels. This hierarchical design enhances fine-grained speaker discrimination while preserving semantic transcription accuracy. Experiments on AliMeeting, AISHELL-4, and AMI-SDM demonstrate that GLSC-SDR achieves competitive or superior performance compared to simulation-based and multi-encoder approaches, without relying on large-scale real conversational data.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25377
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition
Dai, Yuhang
Lin, Haopeng
Qian, Jiale
Yan, Ruiqi
Meng, Hao
Xie, Hanke
Wen, Hanlin
Yin, Shunshun
Tao, Ming
Chen, Xie
Xie, Lei
Wang, Xinsheng
Sound
Large Audio-Language Models (LALMs) have demonstrated remarkable performance in end-to-end speaker diarization and recognition. However, their speaker discriminability remains limited due to the scarcity of large-scale conversational data and the absence of explicit speaker representation optimization. To address this, we propose GLSC-SDR, a paradigm that jointly trains speaker classification with diarization and recognition. We further introduce a Global-Local Speaker Classification strategy, which uses clustered speakers as global labels and re-encoded intra-cluster speakers as local labels. This hierarchical design enhances fine-grained speaker discrimination while preserving semantic transcription accuracy. Experiments on AliMeeting, AISHELL-4, and AMI-SDM demonstrate that GLSC-SDR achieves competitive or superior performance compared to simulation-based and multi-encoder approaches, without relying on large-scale real conversational data.
title Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition
topic Sound
url https://arxiv.org/abs/2603.25377