Read it in Two Steps: Translating Extremely Low-Resource Languages with Code-Augmented Grammar Books

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Chen, Lin, Jiuheng, Liu, Xiao, Zhang, Zekai, Feng, Yansong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916773019779072
author Zhang, Chen
Lin, Jiuheng
Liu, Xiao
Zhang, Zekai
Feng, Yansong
author_facet Zhang, Chen
Lin, Jiuheng
Liu, Xiao
Zhang, Zekai
Feng, Yansong
contents While large language models (LLMs) have shown promise in translating extremely low-resource languages using resources like dictionaries, the effectiveness of grammar books remains debated. This paper investigates the role of grammar books in translating extremely low-resource languages by decomposing it into two key steps: grammar rule retrieval and application. To facilitate the study, we introduce ZhuangRules, a modularized dataset of grammar rules and their corresponding test sentences. Our analysis reveals that rule retrieval constitutes a primary bottleneck in grammar-based translation. Moreover, although LLMs can apply simple rules for translation when explicitly provided, they encounter difficulties in handling more complex rules. To address these challenges, we propose representing grammar rules as code functions, considering their similarities in structure and the benefit of code in facilitating LLM reasoning. Our experiments show that using code rules significantly boosts both rule retrieval and application, ultimately resulting in a 13.1% BLEU improvement in translation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01796
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Read it in Two Steps: Translating Extremely Low-Resource Languages with Code-Augmented Grammar Books
Zhang, Chen
Lin, Jiuheng
Liu, Xiao
Zhang, Zekai
Feng, Yansong
Computation and Language
While large language models (LLMs) have shown promise in translating extremely low-resource languages using resources like dictionaries, the effectiveness of grammar books remains debated. This paper investigates the role of grammar books in translating extremely low-resource languages by decomposing it into two key steps: grammar rule retrieval and application. To facilitate the study, we introduce ZhuangRules, a modularized dataset of grammar rules and their corresponding test sentences. Our analysis reveals that rule retrieval constitutes a primary bottleneck in grammar-based translation. Moreover, although LLMs can apply simple rules for translation when explicitly provided, they encounter difficulties in handling more complex rules. To address these challenges, we propose representing grammar rules as code functions, considering their similarities in structure and the benefit of code in facilitating LLM reasoning. Our experiments show that using code rules significantly boosts both rule retrieval and application, ultimately resulting in a 13.1% BLEU improvement in translation.
title Read it in Two Steps: Translating Extremely Low-Resource Languages with Code-Augmented Grammar Books
topic Computation and Language
url https://arxiv.org/abs/2506.01796