ACADATA: Parallel Dataset of Academic Data for Machine Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lacunza, Iñaki, Gilabert, Javier Garcia, Fornaciari, Francesca De Luca, Aula-Blasco, Javier, Gonzalez-Agirre, Aitor, Melero, Maite, Villegas, Marta
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909846840803328
author Lacunza, Iñaki
Gilabert, Javier Garcia
Fornaciari, Francesca De Luca
Aula-Blasco, Javier
Gonzalez-Agirre, Aitor
Melero, Maite
Villegas, Marta
author_facet Lacunza, Iñaki
Gilabert, Javier Garcia
Fornaciari, Francesca De Luca
Aula-Blasco, Javier
Gonzalez-Agirre, Aitor
Melero, Maite
Villegas, Marta
contents We present ACADATA, a high-quality parallel dataset for academic translation, that consists of two subsets: ACAD-TRAIN, which contains approximately 1.5 million author-generated paragraph pairs across 96 language directions and ACAD-BENCH, a curated evaluation set of almost 6,000 translations covering 12 directions. To validate its utility, we fine-tune two Large Language Models (LLMs) on ACAD-TRAIN and benchmark them on ACAD-BENCH against specialized machine-translation systems, general-purpose, open-weight LLMs, and several large-scale proprietary models. Experimental results demonstrate that fine-tuning on ACAD-TRAIN leads to improvements in academic translation quality by +6.1 and +12.4 d-BLEU points on average for 7B and 2B models respectively, while also improving long-context translation in a general domain by up to 24.9% when translating out of English. The fine-tuned top-performing model surpasses the best propietary and open-weight models on academic translation domain. By releasing ACAD-TRAIN, ACAD-BENCH and the fine-tuned models, we provide the community with a valuable resource to advance research in academic domain and long-context translation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12621
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ACADATA: Parallel Dataset of Academic Data for Machine Translation
Lacunza, Iñaki
Gilabert, Javier Garcia
Fornaciari, Francesca De Luca
Aula-Blasco, Javier
Gonzalez-Agirre, Aitor
Melero, Maite
Villegas, Marta
Computation and Language
We present ACADATA, a high-quality parallel dataset for academic translation, that consists of two subsets: ACAD-TRAIN, which contains approximately 1.5 million author-generated paragraph pairs across 96 language directions and ACAD-BENCH, a curated evaluation set of almost 6,000 translations covering 12 directions. To validate its utility, we fine-tune two Large Language Models (LLMs) on ACAD-TRAIN and benchmark them on ACAD-BENCH against specialized machine-translation systems, general-purpose, open-weight LLMs, and several large-scale proprietary models. Experimental results demonstrate that fine-tuning on ACAD-TRAIN leads to improvements in academic translation quality by +6.1 and +12.4 d-BLEU points on average for 7B and 2B models respectively, while also improving long-context translation in a general domain by up to 24.9% when translating out of English. The fine-tuned top-performing model surpasses the best propietary and open-weight models on academic translation domain. By releasing ACAD-TRAIN, ACAD-BENCH and the fine-tuned models, we provide the community with a valuable resource to advance research in academic domain and long-context translation.
title ACADATA: Parallel Dataset of Academic Data for Machine Translation
topic Computation and Language
url https://arxiv.org/abs/2510.12621