Efficient Low-Resource Language Adaptation via Multi-Source Dynamic Logit Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Chen, Lin, Jiuheng, Liao, Zhiyuan, Feng, Yansong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917422593736704
author Zhang, Chen
Lin, Jiuheng
Liao, Zhiyuan
Feng, Yansong
author_facet Zhang, Chen
Lin, Jiuheng
Liao, Zhiyuan
Feng, Yansong
contents Adapting large language models (LLMs) to low-resource languages (LRLs) is constrained by the scarcity of task data and computational resources. Although Proxy Tuning offers a logit-level strategy for introducing scaling effects, it often fails in LRL settings because the large model's weak LRL competence might overwhelm the knowledge of specialized smaller models. We thus propose TriMix, a test-time logit fusion framework that dynamically balances capabilities from three different sources: LRL competence from a continually pretrained small model, task competence from high-resource language instruction tuning, and the scaling benefits of large models. It is data- and compute-efficient, requiring no LRL task annotations, and only continual pretraining on a small model. Experiments across four model families and eight LRLs show that TriMix consistently outperforms single-model baselines and Proxy Tuning. Our analysis reveals that prioritizing the small LRL-specialized model's logits is crucial for success, challenging the prevalent large-model-dominant assumption.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18106
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Efficient Low-Resource Language Adaptation via Multi-Source Dynamic Logit Fusion
Zhang, Chen
Lin, Jiuheng
Liao, Zhiyuan
Feng, Yansong
Computation and Language
Adapting large language models (LLMs) to low-resource languages (LRLs) is constrained by the scarcity of task data and computational resources. Although Proxy Tuning offers a logit-level strategy for introducing scaling effects, it often fails in LRL settings because the large model's weak LRL competence might overwhelm the knowledge of specialized smaller models. We thus propose TriMix, a test-time logit fusion framework that dynamically balances capabilities from three different sources: LRL competence from a continually pretrained small model, task competence from high-resource language instruction tuning, and the scaling benefits of large models. It is data- and compute-efficient, requiring no LRL task annotations, and only continual pretraining on a small model. Experiments across four model families and eight LRLs show that TriMix consistently outperforms single-model baselines and Proxy Tuning. Our analysis reveals that prioritizing the small LRL-specialized model's logits is crucial for success, challenging the prevalent large-model-dominant assumption.
title Efficient Low-Resource Language Adaptation via Multi-Source Dynamic Logit Fusion
topic Computation and Language
url https://arxiv.org/abs/2604.18106