Corpus of Cross-lingual Dialogues with Minutes and Detection of Misunderstandings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Čechovič, Marko, Komorníková, Natália, Macháček, Dominik, Bojar, Ondřej
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909974595108864
author Čechovič, Marko
Komorníková, Natália
Macháček, Dominik
Bojar, Ondřej
author_facet Čechovič, Marko
Komorníková, Natália
Macháček, Dominik
Bojar, Ondřej
contents Speech processing and translation technology have the potential to facilitate meetings of individuals who do not share any common language. To evaluate automatic systems for such a task, a versatile and realistic evaluation corpus is needed. Therefore, we create and present a corpus of cross-lingual dialogues between individuals without a common language who were facilitated by automatic simultaneous speech translation. The corpus consists of 5 hours of speech recordings with ASR and gold transcripts in 12 original languages and automatic and corrected translations into English. For the purposes of research into cross-lingual summarization, our corpus also includes written summaries (minutes) of the meetings. Moreover, we propose automatic detection of misunderstandings. For an overview of this task and its complexity, we attempt to quantify misunderstandings in cross-lingual meetings. We annotate misunderstandings manually and also test the ability of current large language models to detect them automatically. The results show that the Gemini model is able to identify text spans with misunderstandings with recall of 77% and precision of 47%.
format Preprint
id arxiv_https___arxiv_org_abs_2512_20204
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Corpus of Cross-lingual Dialogues with Minutes and Detection of Misunderstandings
Čechovič, Marko
Komorníková, Natália
Macháček, Dominik
Bojar, Ondřej
Computation and Language
Artificial Intelligence
Speech processing and translation technology have the potential to facilitate meetings of individuals who do not share any common language. To evaluate automatic systems for such a task, a versatile and realistic evaluation corpus is needed. Therefore, we create and present a corpus of cross-lingual dialogues between individuals without a common language who were facilitated by automatic simultaneous speech translation. The corpus consists of 5 hours of speech recordings with ASR and gold transcripts in 12 original languages and automatic and corrected translations into English. For the purposes of research into cross-lingual summarization, our corpus also includes written summaries (minutes) of the meetings. Moreover, we propose automatic detection of misunderstandings. For an overview of this task and its complexity, we attempt to quantify misunderstandings in cross-lingual meetings. We annotate misunderstandings manually and also test the ability of current large language models to detect them automatically. The results show that the Gemini model is able to identify text spans with misunderstandings with recall of 77% and precision of 47%.
title Corpus of Cross-lingual Dialogues with Minutes and Detection of Misunderstandings
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.20204