BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866913893251547136 |
|---|---|
| author | The Omnilingual MT Team Andrews, Pierre Artetxe, Mikel Meglioli, Mariano Coria Costa-jussà, Marta R. Chuang, Joe Dale, David Gao, Cynthia Maillard, Jean Mourachko, Alex Ropers, Christophe Saleem, Safiyyah Sánchez, Eduardo Tsiamas, Ioannis Turkatenko, Arina Ventayol-Boada, Albert Yates, Shireen |
| author_facet | The Omnilingual MT Team Andrews, Pierre Artetxe, Mikel Meglioli, Mariano Coria Costa-jussà, Marta R. Chuang, Joe Dale, David Gao, Cynthia Maillard, Jean Mourachko, Alex Ropers, Christophe Saleem, Safiyyah Sánchez, Eduardo Tsiamas, Ioannis Turkatenko, Arina Ventayol-Boada, Albert Yates, Shireen |
| contents | BOUQuET is a multi-way, multicentric and multi-register/domain dataset and benchmark, and a broader collaborative initiative. This dataset is handcrafted in 8 non-English languages. Each of these source languages are representative of the most widely spoken ones and therefore they have the potential to serve as pivot languages that will enable more accurate translations. The dataset is multicentric to enforce representation of multilingual language features. In addition, the dataset goes beyond the sentence level, as it is organized in paragraphs of various lengths. Compared with related machine translation datasets, we show that BOUQuET has a broader representation of domains while simplifying the translation task for non-experts. Therefore, BOUQuET is specially suitable for crowd-source extension for which we are launching a call aiming at collecting a multi-way parallel corpus covering any written language. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_04314 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation The Omnilingual MT Team Andrews, Pierre Artetxe, Mikel Meglioli, Mariano Coria Costa-jussà, Marta R. Chuang, Joe Dale, David Gao, Cynthia Maillard, Jean Mourachko, Alex Ropers, Christophe Saleem, Safiyyah Sánchez, Eduardo Tsiamas, Ioannis Turkatenko, Arina Ventayol-Boada, Albert Yates, Shireen Computation and Language I.2.7 BOUQuET is a multi-way, multicentric and multi-register/domain dataset and benchmark, and a broader collaborative initiative. This dataset is handcrafted in 8 non-English languages. Each of these source languages are representative of the most widely spoken ones and therefore they have the potential to serve as pivot languages that will enable more accurate translations. The dataset is multicentric to enforce representation of multilingual language features. In addition, the dataset goes beyond the sentence level, as it is organized in paragraphs of various lengths. Compared with related machine translation datasets, we show that BOUQuET has a broader representation of domains while simplifying the translation task for non-experts. Therefore, BOUQuET is specially suitable for crowd-source extension for which we are launching a call aiming at collecting a multi-way parallel corpus covering any written language. |
| title | BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation |
| topic | Computation and Language I.2.7 |
| url | https://arxiv.org/abs/2502.04314 |