Typologically-Informed Candidate Reranking for LLM-based Translation into Low-Resource Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abeykoon, Nipuna, Weerathunga, Ashen, Wijesinghe, Pubudu, Krishnamurthy, Parameswari
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917239529144320
author Abeykoon, Nipuna
Weerathunga, Ashen
Wijesinghe, Pubudu
Krishnamurthy, Parameswari
author_facet Abeykoon, Nipuna
Weerathunga, Ashen
Wijesinghe, Pubudu
Krishnamurthy, Parameswari
contents Large language models trained predominantly on high-resource languages exhibit systematic biases toward dominant typological patterns, leading to structural non-conformance when translating into typologically divergent low-resource languages. We present a framework that leverages linguistic typology to improve translation quality without parallel training data or model retraining. The framework consists of two components: the Universal Metalinguistic Framework (UMF), which represents languages as structured profiles across 16 typological dimensions with divergence-weighted scoring, and the Computational Engine, which operates through linguistic disambiguation during generation and typological compliance scoring during selection. Evaluation across nine language pairs demonstrates intervention rates strongly correlating with typological distance from English. In experiments on 341 English sentences each having different morphological and syntactic phenomena, the framework shows an intervention precision of 48.16% for conservatively treated languages, 28.15% for morphologically dense languages, and 86.26% for structurally profiled languages. The framework requires no parallel training data and operates with any LLM capable of producing multiple candidate outputs, enabling practical deployment for under-resourced languages.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01162
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Typologically-Informed Candidate Reranking for LLM-based Translation into Low-Resource Languages
Abeykoon, Nipuna
Weerathunga, Ashen
Wijesinghe, Pubudu
Krishnamurthy, Parameswari
Computation and Language
Large language models trained predominantly on high-resource languages exhibit systematic biases toward dominant typological patterns, leading to structural non-conformance when translating into typologically divergent low-resource languages. We present a framework that leverages linguistic typology to improve translation quality without parallel training data or model retraining. The framework consists of two components: the Universal Metalinguistic Framework (UMF), which represents languages as structured profiles across 16 typological dimensions with divergence-weighted scoring, and the Computational Engine, which operates through linguistic disambiguation during generation and typological compliance scoring during selection. Evaluation across nine language pairs demonstrates intervention rates strongly correlating with typological distance from English. In experiments on 341 English sentences each having different morphological and syntactic phenomena, the framework shows an intervention precision of 48.16% for conservatively treated languages, 28.15% for morphologically dense languages, and 86.26% for structurally profiled languages. The framework requires no parallel training data and operates with any LLM capable of producing multiple candidate outputs, enabling practical deployment for under-resourced languages.
title Typologically-Informed Candidate Reranking for LLM-based Translation into Low-Resource Languages
topic Computation and Language
url https://arxiv.org/abs/2602.01162