Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Medennikov, Ivan, Park, Taejin, Wang, Weiqing, Huang, He, Dhawan, Kunal, Wang, Jinhan, Balam, Jagadeesh, Ginsburg, Boris
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911074989637632
author Medennikov, Ivan
Park, Taejin
Wang, Weiqing
Huang, He
Dhawan, Kunal
Wang, Jinhan
Balam, Jagadeesh
Ginsburg, Boris
author_facet Medennikov, Ivan
Park, Taejin
Wang, Weiqing
Huang, He
Dhawan, Kunal
Wang, Jinhan
Balam, Jagadeesh
Ginsburg, Boris
contents This paper presents a streaming extension for the Sortformer speaker diarization framework, whose key property is the arrival-time ordering of output speakers. The proposed approach employs an Arrival-Order Speaker Cache (AOSC) to store frame-level acoustic embeddings of previously observed speakers. Unlike conventional speaker-tracing buffers, AOSC orders embeddings by speaker index corresponding to their arrival time order, and is dynamically updated by selecting frames with the highest scores based on the model's past predictions. Notably, the number of stored embeddings per speaker is determined dynamically by the update mechanism, ensuring efficient cache utilization and precise speaker tracking. Experiments on benchmark datasets confirm the effectiveness and flexibility of our approach, even in low-latency setups. These results establish Streaming Sortformer as a robust solution for real-time multi-speaker tracking and a foundation for streaming multi-talker speech processing.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18446
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
Medennikov, Ivan
Park, Taejin
Wang, Weiqing
Huang, He
Dhawan, Kunal
Wang, Jinhan
Balam, Jagadeesh
Ginsburg, Boris
Audio and Speech Processing
Sound
This paper presents a streaming extension for the Sortformer speaker diarization framework, whose key property is the arrival-time ordering of output speakers. The proposed approach employs an Arrival-Order Speaker Cache (AOSC) to store frame-level acoustic embeddings of previously observed speakers. Unlike conventional speaker-tracing buffers, AOSC orders embeddings by speaker index corresponding to their arrival time order, and is dynamically updated by selecting frames with the highest scores based on the model's past predictions. Notably, the number of stored embeddings per speaker is determined dynamically by the update mechanism, ensuring efficient cache utilization and precise speaker tracking. Experiments on benchmark datasets confirm the effectiveness and flexibility of our approach, even in low-latency setups. These results establish Streaming Sortformer as a robust solution for real-time multi-speaker tracking and a foundation for streaming multi-talker speech processing.
title Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2507.18446