Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918412538609664 |
|---|---|
| author | Dai, Yuhang Lin, Haopeng Qian, Jiale Yan, Ruiqi Meng, Hao Xie, Hanke Wen, Hanlin Yin, Shunshun Tao, Ming Chen, Xie Xie, Lei Wang, Xinsheng |
| author_facet | Dai, Yuhang Lin, Haopeng Qian, Jiale Yan, Ruiqi Meng, Hao Xie, Hanke Wen, Hanlin Yin, Shunshun Tao, Ming Chen, Xie Xie, Lei Wang, Xinsheng |
| contents | Large Audio-Language Models (LALMs) have demonstrated remarkable performance in end-to-end speaker diarization and recognition. However, their speaker discriminability remains limited due to the scarcity of large-scale conversational data and the absence of explicit speaker representation optimization. To address this, we propose GLSC-SDR, a paradigm that jointly trains speaker classification with diarization and recognition. We further introduce a Global-Local Speaker Classification strategy, which uses clustered speakers as global labels and re-encoded intra-cluster speakers as local labels. This hierarchical design enhances fine-grained speaker discrimination while preserving semantic transcription accuracy. Experiments on AliMeeting, AISHELL-4, and AMI-SDM demonstrate that GLSC-SDR achieves competitive or superior performance compared to simulation-based and multi-encoder approaches, without relying on large-scale real conversational data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_25377 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition Dai, Yuhang Lin, Haopeng Qian, Jiale Yan, Ruiqi Meng, Hao Xie, Hanke Wen, Hanlin Yin, Shunshun Tao, Ming Chen, Xie Xie, Lei Wang, Xinsheng Sound Large Audio-Language Models (LALMs) have demonstrated remarkable performance in end-to-end speaker diarization and recognition. However, their speaker discriminability remains limited due to the scarcity of large-scale conversational data and the absence of explicit speaker representation optimization. To address this, we propose GLSC-SDR, a paradigm that jointly trains speaker classification with diarization and recognition. We further introduce a Global-Local Speaker Classification strategy, which uses clustered speakers as global labels and re-encoded intra-cluster speakers as local labels. This hierarchical design enhances fine-grained speaker discrimination while preserving semantic transcription accuracy. Experiments on AliMeeting, AISHELL-4, and AMI-SDM demonstrate that GLSC-SDR achieves competitive or superior performance compared to simulation-based and multi-encoder approaches, without relying on large-scale real conversational data. |
| title | Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition |
| topic | Sound |
| url | https://arxiv.org/abs/2603.25377 |