ACADATA: Parallel Dataset of Academic Data for Machine Translation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909846840803328 |
|---|---|
| author | Lacunza, Iñaki Gilabert, Javier Garcia Fornaciari, Francesca De Luca Aula-Blasco, Javier Gonzalez-Agirre, Aitor Melero, Maite Villegas, Marta |
| author_facet | Lacunza, Iñaki Gilabert, Javier Garcia Fornaciari, Francesca De Luca Aula-Blasco, Javier Gonzalez-Agirre, Aitor Melero, Maite Villegas, Marta |
| contents | We present ACADATA, a high-quality parallel dataset for academic translation, that consists of two subsets: ACAD-TRAIN, which contains approximately 1.5 million author-generated paragraph pairs across 96 language directions and ACAD-BENCH, a curated evaluation set of almost 6,000 translations covering 12 directions. To validate its utility, we fine-tune two Large Language Models (LLMs) on ACAD-TRAIN and benchmark them on ACAD-BENCH against specialized machine-translation systems, general-purpose, open-weight LLMs, and several large-scale proprietary models. Experimental results demonstrate that fine-tuning on ACAD-TRAIN leads to improvements in academic translation quality by +6.1 and +12.4 d-BLEU points on average for 7B and 2B models respectively, while also improving long-context translation in a general domain by up to 24.9% when translating out of English. The fine-tuned top-performing model surpasses the best propietary and open-weight models on academic translation domain. By releasing ACAD-TRAIN, ACAD-BENCH and the fine-tuned models, we provide the community with a valuable resource to advance research in academic domain and long-context translation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_12621 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ACADATA: Parallel Dataset of Academic Data for Machine Translation Lacunza, Iñaki Gilabert, Javier Garcia Fornaciari, Francesca De Luca Aula-Blasco, Javier Gonzalez-Agirre, Aitor Melero, Maite Villegas, Marta Computation and Language We present ACADATA, a high-quality parallel dataset for academic translation, that consists of two subsets: ACAD-TRAIN, which contains approximately 1.5 million author-generated paragraph pairs across 96 language directions and ACAD-BENCH, a curated evaluation set of almost 6,000 translations covering 12 directions. To validate its utility, we fine-tune two Large Language Models (LLMs) on ACAD-TRAIN and benchmark them on ACAD-BENCH against specialized machine-translation systems, general-purpose, open-weight LLMs, and several large-scale proprietary models. Experimental results demonstrate that fine-tuning on ACAD-TRAIN leads to improvements in academic translation quality by +6.1 and +12.4 d-BLEU points on average for 7B and 2B models respectively, while also improving long-context translation in a general domain by up to 24.9% when translating out of English. The fine-tuned top-performing model surpasses the best propietary and open-weight models on academic translation domain. By releasing ACAD-TRAIN, ACAD-BENCH and the fine-tuned models, we provide the community with a valuable resource to advance research in academic domain and long-context translation. |
| title | ACADATA: Parallel Dataset of Academic Data for Machine Translation |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.12621 |