MOVER: Combining Multiple Meeting Recognition Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kamo, Naoyuki, Ochiai, Tsubasa, Delcroix, Marc, Nakatani, Tomohiro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913978817445888
author Kamo, Naoyuki
Ochiai, Tsubasa
Delcroix, Marc
Nakatani, Tomohiro
author_facet Kamo, Naoyuki
Ochiai, Tsubasa
Delcroix, Marc
Nakatani, Tomohiro
contents In this paper, we propose Meeting recognizer Output Voting Error Reduction (MOVER), a novel system combination method for meeting recognition tasks. Although there are methods to combine the output of diarization (e.g., DOVER) or automatic speech recognition (ASR) systems (e.g., ROVER), MOVER is the first approach that can combine the outputs of meeting recognition systems that differ in terms of both diarization and ASR. MOVER combines hypotheses with different time intervals and speaker labels through a five-stage process that includes speaker alignment, segment grouping, word and timing combination, etc. Experimental results on the CHiME-8 DASR task and the multi-channel track of the NOTSOFAR-1 task demonstrate that MOVER can successfully combine multiple meeting recognition systems with diverse diarization and recognition outputs, achieving relative tcpWER improvements of 9.55 % and 8.51 % over the state-of-the-art systems for both tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05055
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MOVER: Combining Multiple Meeting Recognition Systems
Kamo, Naoyuki
Ochiai, Tsubasa
Delcroix, Marc
Nakatani, Tomohiro
Audio and Speech Processing
In this paper, we propose Meeting recognizer Output Voting Error Reduction (MOVER), a novel system combination method for meeting recognition tasks. Although there are methods to combine the output of diarization (e.g., DOVER) or automatic speech recognition (ASR) systems (e.g., ROVER), MOVER is the first approach that can combine the outputs of meeting recognition systems that differ in terms of both diarization and ASR. MOVER combines hypotheses with different time intervals and speaker labels through a five-stage process that includes speaker alignment, segment grouping, word and timing combination, etc. Experimental results on the CHiME-8 DASR task and the multi-channel track of the NOTSOFAR-1 task demonstrate that MOVER can successfully combine multiple meeting recognition systems with diverse diarization and recognition outputs, achieving relative tcpWER improvements of 9.55 % and 8.51 % over the state-of-the-art systems for both tasks.
title MOVER: Combining Multiple Meeting Recognition Systems
topic Audio and Speech Processing
url https://arxiv.org/abs/2508.05055