Low-resource Machine Translation: what for? who for? An observational study on a dedicated Tetun language translation service

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Merx, Raphael, Correia, Adérito José Guterres, Suominen, Hanna, Vylomova, Ekaterina
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910993662083072
author Merx, Raphael
Correia, Adérito José Guterres
Suominen, Hanna
Vylomova, Ekaterina
author_facet Merx, Raphael
Correia, Adérito José Guterres
Suominen, Hanna
Vylomova, Ekaterina
contents Low-resource machine translation (MT) presents a diversity of community needs and application challenges that remain poorly understood. To complement surveys and focus groups, which tend to rely on small samples of respondents, we propose an observational study on actual usage patterns of tetun$.$org, a specialized MT service for the Tetun language, which is the lingua franca in Timor-Leste. Our analysis of 100,000 translation requests reveals patterns that challenge assumptions based on existing corpora. We find that users, many of them students on mobile devices, typically translate text from a high-resource language into Tetun across diverse domains including science, healthcare, and daily life. This contrasts sharply with available Tetun corpora, which are dominated by news articles covering government and social issues. Our results suggest that MT systems for institutionalized minority languages like Tetun should prioritize accuracy on domains relevant to educational contexts, in the high-resource to low-resource direction. More broadly, this study demonstrates how observational analysis can inform low-resource language technology development, by grounding research in practical community needs.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12262
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Low-resource Machine Translation: what for? who for? An observational study on a dedicated Tetun language translation service
Merx, Raphael
Correia, Adérito José Guterres
Suominen, Hanna
Vylomova, Ekaterina
Computation and Language
Artificial Intelligence
Low-resource machine translation (MT) presents a diversity of community needs and application challenges that remain poorly understood. To complement surveys and focus groups, which tend to rely on small samples of respondents, we propose an observational study on actual usage patterns of tetun$.$org, a specialized MT service for the Tetun language, which is the lingua franca in Timor-Leste. Our analysis of 100,000 translation requests reveals patterns that challenge assumptions based on existing corpora. We find that users, many of them students on mobile devices, typically translate text from a high-resource language into Tetun across diverse domains including science, healthcare, and daily life. This contrasts sharply with available Tetun corpora, which are dominated by news articles covering government and social issues. Our results suggest that MT systems for institutionalized minority languages like Tetun should prioritize accuracy on domains relevant to educational contexts, in the high-resource to low-resource direction. More broadly, this study demonstrates how observational analysis can inform low-resource language technology development, by grounding research in practical community needs.
title Low-resource Machine Translation: what for? who for? An observational study on a dedicated Tetun language translation service
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.12262