A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nguyen, Thai-Binh, Zmolikova, Katerina, Ma, Pingchuan, Pham, Ngoc Quan, Fuegen, Christian, Waibel, Alexander
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912898410872832
author Nguyen, Thai-Binh
Zmolikova, Katerina
Ma, Pingchuan
Pham, Ngoc Quan
Fuegen, Christian
Waibel, Alexander
author_facet Nguyen, Thai-Binh
Zmolikova, Katerina
Ma, Pingchuan
Pham, Ngoc Quan
Fuegen, Christian
Waibel, Alexander
contents We introduce the task of Multi-Modal Context-Aware Recognition (MCoRec) in the ninth CHiME Challenge, which addresses the cocktail-party problem of overlapping conversations in a single-room setting using audio, visual, and contextual cues. MCoRec captures natural multi-party conversations where the recordings focus on unscripted, casual group chats, leading to extreme speech overlap of up to 100% and highly fragmented conversational turns. The task requires systems to answer the question "Who speaks when, what, and with whom?" by jointly transcribing each speaker's speech and clustering them into their respective conversations from audio-visual recordings. Audio-only baselines exceed 100% word error rate, whereas incorporating visual cues yields substantial 50% improvements, highlighting the importance of multi-modality. In this manuscript, we present the motivation behind the task, outline the data collection process, and report the baseline systems developed for the MCoRec.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23276
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results
Nguyen, Thai-Binh
Zmolikova, Katerina
Ma, Pingchuan
Pham, Ngoc Quan
Fuegen, Christian
Waibel, Alexander
Computation and Language
We introduce the task of Multi-Modal Context-Aware Recognition (MCoRec) in the ninth CHiME Challenge, which addresses the cocktail-party problem of overlapping conversations in a single-room setting using audio, visual, and contextual cues. MCoRec captures natural multi-party conversations where the recordings focus on unscripted, casual group chats, leading to extreme speech overlap of up to 100% and highly fragmented conversational turns. The task requires systems to answer the question "Who speaks when, what, and with whom?" by jointly transcribing each speaker's speech and clustering them into their respective conversations from audio-visual recordings. Audio-only baselines exceed 100% word error rate, whereas incorporating visual cues yields substantial 50% improvements, highlighting the importance of multi-modality. In this manuscript, we present the motivation behind the task, outline the data collection process, and report the baseline systems developed for the MCoRec.
title A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results
topic Computation and Language
url https://arxiv.org/abs/2510.23276