MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929224667889664 |
|---|---|
| author | von Neumann, Thilo Boeddeker, Christoph Delcroix, Marc Haeb-Umbach, Reinhold |
| author_facet | von Neumann, Thilo Boeddeker, Christoph Delcroix, Marc Haeb-Umbach, Reinhold |
| contents | MeetEval is an open-source toolkit to evaluate all kinds of meeting transcription systems. It provides a unified interface for the computation of commonly used Word Error Rates (WERs), specifically cpWER, ORC-WER and MIMO-WER along other WER definitions. We extend the cpWER computation by a temporal constraint to ensure that only words are identified as correct when the temporal alignment is plausible. This leads to a better quality of the matching of the hypothesis string to the reference string that more closely resembles the actual transcription quality, and a system is penalized if it provides poor time annotations. Since word-level timing information is often not available, we present a way to approximate exact word-level timings from segment-level timings (e.g., a sentence) and show that the approximation leads to a similar WER as a matching with exact word-level annotations. At the same time, the time constraint leads to a speedup of the matching algorithm, which outweighs the additional overhead caused by processing the time stamps. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2307_11394 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems von Neumann, Thilo Boeddeker, Christoph Delcroix, Marc Haeb-Umbach, Reinhold Computation and Language Audio and Speech Processing MeetEval is an open-source toolkit to evaluate all kinds of meeting transcription systems. It provides a unified interface for the computation of commonly used Word Error Rates (WERs), specifically cpWER, ORC-WER and MIMO-WER along other WER definitions. We extend the cpWER computation by a temporal constraint to ensure that only words are identified as correct when the temporal alignment is plausible. This leads to a better quality of the matching of the hypothesis string to the reference string that more closely resembles the actual transcription quality, and a system is penalized if it provides poor time annotations. Since word-level timing information is often not available, we present a way to approximate exact word-level timings from segment-level timings (e.g., a sentence) and show that the approximation leads to a similar WER as a matching with exact word-level annotations. At the same time, the time constraint leads to a speedup of the matching algorithm, which outweighs the additional overhead caused by processing the time stamps. |
| title | MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems |
| topic | Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2307.11394 |