Synthetic Function Demonstrations Improve Generation in Low-Resource Programming Languages

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: McKenna, Nick, Xu, Xinnuo, Williams, Jack, Wilson, Nick, Van Durme, Benjamin, Poelitz, Christian
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913113044942848
author McKenna, Nick
Xu, Xinnuo
Williams, Jack
Wilson, Nick
Van Durme, Benjamin
Poelitz, Christian
author_facet McKenna, Nick
Xu, Xinnuo
Williams, Jack
Wilson, Nick
Van Durme, Benjamin
Poelitz, Christian
contents A key consideration when training an LLM is whether the target language is more or less resourced, for example English compared to Welsh, or Python compared to Excel. Typical training data for programming languages consists of real program demonstrations coupled with explanatory human-written comments. In this work we present a novel approach to the creation of such data for low resource programming languages, which lack naturally occurring data. Our process generates synthetic, textbook-quality demonstrations of how to use library functions, which we show makes for good model finetuning data. We demonstrate in an example domain of Excel Formulas. First, we collate language documentation, then we use this to augment a powerful teacher model which generates synthetic training data, and finally finetune student models on the demonstrations. Our technique improves student performance on 2 question-answering datasets: WikiTQ and TAT-QA. We also show advantages of finetuning over standard RAG approaches, which can offer only modest improvement due to the unfamiliarity of the target domain to student models.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18760
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Synthetic Function Demonstrations Improve Generation in Low-Resource Programming Languages
McKenna, Nick
Xu, Xinnuo
Williams, Jack
Wilson, Nick
Van Durme, Benjamin
Poelitz, Christian
Computation and Language
A key consideration when training an LLM is whether the target language is more or less resourced, for example English compared to Welsh, or Python compared to Excel. Typical training data for programming languages consists of real program demonstrations coupled with explanatory human-written comments. In this work we present a novel approach to the creation of such data for low resource programming languages, which lack naturally occurring data. Our process generates synthetic, textbook-quality demonstrations of how to use library functions, which we show makes for good model finetuning data. We demonstrate in an example domain of Excel Formulas. First, we collate language documentation, then we use this to augment a powerful teacher model which generates synthetic training data, and finally finetune student models on the demonstrations. Our technique improves student performance on 2 question-answering datasets: WikiTQ and TAT-QA. We also show advantages of finetuning over standard RAG approaches, which can offer only modest improvement due to the unfamiliarity of the target domain to student models.
title Synthetic Function Demonstrations Improve Generation in Low-Resource Programming Languages
topic Computation and Language
url https://arxiv.org/abs/2503.18760