CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Liangbin, Liao, Xiaohua, Cui, Chaoqun, Wang, Shijing, Huang, Zhaolong, Du, Yanlong, Mao, Wenji
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910057398009856
author Huang, Liangbin
Liao, Xiaohua
Cui, Chaoqun
Wang, Shijing
Huang, Zhaolong
Du, Yanlong
Mao, Wenji
author_facet Huang, Liangbin
Liao, Xiaohua
Cui, Chaoqun
Wang, Shijing
Huang, Zhaolong
Du, Yanlong
Mao, Wenji
contents Traditional speaker diarization systems have primarily focused on constrained scenarios such as meetings and interviews, where the number of speakers is limited and acoustic conditions are relatively clean. To explore open-world speaker diarization, we extend this task to the visual media domain, encompassing complex audiovisual programs such as films and TV series. This new setting introduces several challenges, including long-form video understanding, a large number of speakers, cross-modal asynchrony between audio and visual cues, and uncontrolled in-the-wild variability. To address these challenges, we propose Cinematic Speaker Registration & Diarization (CineSRD), a unified multimodal framework that leverages visual, acoustic, and linguistic cues from video, speech, and subtitles for speaker annotation. CineSRD first performs visual anchor clustering to register initial speakers and then integrates an audio language model for speaker turn detection, refining annotations and supplementing unregistered off-screen speakers. Furthermore, we construct and release a dedicated speaker diarization benchmark for visual media that includes Chinese and English programs. Experimental results demonstrate that CineSRD achieves superior performance on the proposed benchmark and competitive results on conventional datasets, validating its robustness and generalizability in open-world visual media settings.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16966
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization
Huang, Liangbin
Liao, Xiaohua
Cui, Chaoqun
Wang, Shijing
Huang, Zhaolong
Du, Yanlong
Mao, Wenji
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Sound
Audio and Speech Processing
Traditional speaker diarization systems have primarily focused on constrained scenarios such as meetings and interviews, where the number of speakers is limited and acoustic conditions are relatively clean. To explore open-world speaker diarization, we extend this task to the visual media domain, encompassing complex audiovisual programs such as films and TV series. This new setting introduces several challenges, including long-form video understanding, a large number of speakers, cross-modal asynchrony between audio and visual cues, and uncontrolled in-the-wild variability. To address these challenges, we propose Cinematic Speaker Registration & Diarization (CineSRD), a unified multimodal framework that leverages visual, acoustic, and linguistic cues from video, speech, and subtitles for speaker annotation. CineSRD first performs visual anchor clustering to register initial speakers and then integrates an audio language model for speaker turn detection, refining annotations and supplementing unregistered off-screen speakers. Furthermore, we construct and release a dedicated speaker diarization benchmark for visual media that includes Chinese and English programs. Experimental results demonstrate that CineSRD achieves superior performance on the proposed benchmark and competitive results on conventional datasets, validating its robustness and generalizability in open-world visual media settings.
title CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2603.16966