Cross-lingual Transfer in Programming Languages: An Extensive Empirical Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baltaji, Razan, Pujar, Saurabh, Mandel, Louis, Hirzel, Martin, Buratti, Luca, Varshney, Lav
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910996694564864
author Baltaji, Razan
Pujar, Saurabh
Mandel, Louis
Hirzel, Martin
Buratti, Luca
Varshney, Lav
author_facet Baltaji, Razan
Pujar, Saurabh
Mandel, Louis
Hirzel, Martin
Buratti, Luca
Varshney, Lav
contents Large language models (LLMs) have achieved state-of-the-art performance in various software engineering tasks, including error detection, clone detection, and code translation, primarily leveraging high-resource programming languages like Python and Java. However, many critical languages, such as COBOL, as well as emerging languages, such as Rust and Swift, remain low-resource due to limited openly available code. This scarcity hampers the training and effectiveness of LLMs for these languages, increasing software maintenance costs and stifling innovation. Addressing this gap, we investigate the potential of transfer learning to enhance LLM performance on low-resource programming languages by leveraging data from high-resource counterparts. Our extensive empirical study evaluates transferability across 10 to 41 programming languages and five key tasks: code generation, clone detection, code repair, solution domain classification, and error detection. Additionally, we develop a performance prediction model to guess the best source languages for a given target and task, and analyze the features that influence transfer performance. We further replicate a representative subset of experiments with a larger model to test the generalizability of our conclusions to contemporary large-scale LLMs. Our findings demonstrate that cross-lingual transfer significantly outperforms zero-shot learning, with effectiveness varying based on both source and target languages. Furthermore, our model reliably predicts successful transfer sources by considering linguistic and dataset-specific features, offering practical guidance for data acquisition and model training. This work contributes to the development of LLM-driven tools for low-resource programming languages and provides insights into the characteristics that facilitate transfer across language pairs.
format Preprint
id arxiv_https___arxiv_org_abs_2310_16937
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Cross-lingual Transfer in Programming Languages: An Extensive Empirical Study
Baltaji, Razan
Pujar, Saurabh
Mandel, Louis
Hirzel, Martin
Buratti, Luca
Varshney, Lav
Computation and Language
I.2.7; I.2.5
Large language models (LLMs) have achieved state-of-the-art performance in various software engineering tasks, including error detection, clone detection, and code translation, primarily leveraging high-resource programming languages like Python and Java. However, many critical languages, such as COBOL, as well as emerging languages, such as Rust and Swift, remain low-resource due to limited openly available code. This scarcity hampers the training and effectiveness of LLMs for these languages, increasing software maintenance costs and stifling innovation. Addressing this gap, we investigate the potential of transfer learning to enhance LLM performance on low-resource programming languages by leveraging data from high-resource counterparts. Our extensive empirical study evaluates transferability across 10 to 41 programming languages and five key tasks: code generation, clone detection, code repair, solution domain classification, and error detection. Additionally, we develop a performance prediction model to guess the best source languages for a given target and task, and analyze the features that influence transfer performance. We further replicate a representative subset of experiments with a larger model to test the generalizability of our conclusions to contemporary large-scale LLMs. Our findings demonstrate that cross-lingual transfer significantly outperforms zero-shot learning, with effectiveness varying based on both source and target languages. Furthermore, our model reliably predicts successful transfer sources by considering linguistic and dataset-specific features, offering practical guidance for data acquisition and model training. This work contributes to the development of LLM-driven tools for low-resource programming languages and provides insights into the characteristics that facilitate transfer across language pairs.
title Cross-lingual Transfer in Programming Languages: An Extensive Empirical Study
topic Computation and Language
I.2.7; I.2.5
url https://arxiv.org/abs/2310.16937