Machine Translation Models are Zero-Shot Detectors of Translation Direction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wastl, Michelle, Vamvas, Jannis, Sennrich, Rico
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913862550290432
author Wastl, Michelle
Vamvas, Jannis
Sennrich, Rico
author_facet Wastl, Michelle
Vamvas, Jannis
Sennrich, Rico
contents Detecting the translation direction of parallel text has applications for machine translation training and evaluation, but also has forensic applications such as resolving plagiarism or forgery allegations. In this work, we explore an unsupervised approach to translation direction detection based on the simple hypothesis that $p(\text{translation}|\text{original})>p(\text{original}|\text{translation})$, motivated by the well-known simplification effect in translationese or machine-translationese. In experiments with massively multilingual machine translation models across 20 translation directions, we confirm the effectiveness of the approach for high-resource language pairs, achieving document-level accuracies of 82--96% for NMT-produced translations, and 60--81% for human translations, depending on the model used. Code and demo are available at https://github.com/ZurichNLP/translation-direction-detection
format Preprint
id arxiv_https___arxiv_org_abs_2401_06769
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Machine Translation Models are Zero-Shot Detectors of Translation Direction
Wastl, Michelle
Vamvas, Jannis
Sennrich, Rico
Computation and Language
Detecting the translation direction of parallel text has applications for machine translation training and evaluation, but also has forensic applications such as resolving plagiarism or forgery allegations. In this work, we explore an unsupervised approach to translation direction detection based on the simple hypothesis that $p(\text{translation}|\text{original})>p(\text{original}|\text{translation})$, motivated by the well-known simplification effect in translationese or machine-translationese. In experiments with massively multilingual machine translation models across 20 translation directions, we confirm the effectiveness of the approach for high-resource language pairs, achieving document-level accuracies of 82--96% for NMT-produced translations, and 60--81% for human translations, depending on the model used. Code and demo are available at https://github.com/ZurichNLP/translation-direction-detection
title Machine Translation Models are Zero-Shot Detectors of Translation Direction
topic Computation and Language
url https://arxiv.org/abs/2401.06769