BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: The Omnilingual MT Team, Andrews, Pierre, Artetxe, Mikel, Meglioli, Mariano Coria, Costa-jussà, Marta R., Chuang, Joe, Dale, David, Gao, Cynthia, Maillard, Jean, Mourachko, Alex, Ropers, Christophe, Saleem, Safiyyah, Sánchez, Eduardo, Tsiamas, Ioannis, Turkatenko, Arina, Ventayol-Boada, Albert, Yates, Shireen
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913893251547136
author The Omnilingual MT Team
Andrews, Pierre
Artetxe, Mikel
Meglioli, Mariano Coria
Costa-jussà, Marta R.
Chuang, Joe
Dale, David
Gao, Cynthia
Maillard, Jean
Mourachko, Alex
Ropers, Christophe
Saleem, Safiyyah
Sánchez, Eduardo
Tsiamas, Ioannis
Turkatenko, Arina
Ventayol-Boada, Albert
Yates, Shireen
author_facet The Omnilingual MT Team
Andrews, Pierre
Artetxe, Mikel
Meglioli, Mariano Coria
Costa-jussà, Marta R.
Chuang, Joe
Dale, David
Gao, Cynthia
Maillard, Jean
Mourachko, Alex
Ropers, Christophe
Saleem, Safiyyah
Sánchez, Eduardo
Tsiamas, Ioannis
Turkatenko, Arina
Ventayol-Boada, Albert
Yates, Shireen
contents BOUQuET is a multi-way, multicentric and multi-register/domain dataset and benchmark, and a broader collaborative initiative. This dataset is handcrafted in 8 non-English languages. Each of these source languages are representative of the most widely spoken ones and therefore they have the potential to serve as pivot languages that will enable more accurate translations. The dataset is multicentric to enforce representation of multilingual language features. In addition, the dataset goes beyond the sentence level, as it is organized in paragraphs of various lengths. Compared with related machine translation datasets, we show that BOUQuET has a broader representation of domains while simplifying the translation task for non-experts. Therefore, BOUQuET is specially suitable for crowd-source extension for which we are launching a call aiming at collecting a multi-way parallel corpus covering any written language.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04314
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
The Omnilingual MT Team
Andrews, Pierre
Artetxe, Mikel
Meglioli, Mariano Coria
Costa-jussà, Marta R.
Chuang, Joe
Dale, David
Gao, Cynthia
Maillard, Jean
Mourachko, Alex
Ropers, Christophe
Saleem, Safiyyah
Sánchez, Eduardo
Tsiamas, Ioannis
Turkatenko, Arina
Ventayol-Boada, Albert
Yates, Shireen
Computation and Language
I.2.7
BOUQuET is a multi-way, multicentric and multi-register/domain dataset and benchmark, and a broader collaborative initiative. This dataset is handcrafted in 8 non-English languages. Each of these source languages are representative of the most widely spoken ones and therefore they have the potential to serve as pivot languages that will enable more accurate translations. The dataset is multicentric to enforce representation of multilingual language features. In addition, the dataset goes beyond the sentence level, as it is organized in paragraphs of various lengths. Compared with related machine translation datasets, we show that BOUQuET has a broader representation of domains while simplifying the translation task for non-experts. Therefore, BOUQuET is specially suitable for crowd-source extension for which we are launching a call aiming at collecting a multi-way parallel corpus covering any written language.
title BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2502.04314