Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Uchiyama, Fumiya, Kojima, Takeshi, Gambardella, Andrew, Cao, Qi, Iwasawa, Yusuke, Matsuo, Yutaka
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909663779356672
author Uchiyama, Fumiya
Kojima, Takeshi
Gambardella, Andrew
Cao, Qi
Iwasawa, Yusuke
Matsuo, Yutaka
author_facet Uchiyama, Fumiya
Kojima, Takeshi
Gambardella, Andrew
Cao, Qi
Iwasawa, Yusuke
Matsuo, Yutaka
contents Recent large language models (LLMs) have demonstrated remarkable generalization abilities in mathematics and logical reasoning tasks. Prior research indicates that LLMs pre-trained with programming language data exhibit high mathematical and reasoning abilities; however, this causal relationship has not been rigorously tested. Our research aims to verify which programming languages and features during pre-training affect logical inference performance. Specifically, we pre-trained decoder-based language models from scratch using datasets from ten programming languages (e.g., Python, C, Java) and three natural language datasets (Wikipedia, Fineweb, C4) under identical conditions. Thereafter, we evaluated the trained models in a few-shot in-context learning setting on logical reasoning tasks: FLD and bAbi, which do not require commonsense or world knowledge. The results demonstrate that nearly all models trained with programming languages consistently outperform those trained with natural languages, indicating that programming languages contain factors that elicit logic inference performance. In addition, we found that models trained with programming languages exhibit a better ability to follow instructions compared to those trained with natural languages. Further analysis reveals that the depth of Abstract Syntax Trees representing parsed results of programs also affects logical reasoning performance. These findings will offer insights into the essential elements of pre-training for acquiring the foundational abilities of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06735
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance?
Uchiyama, Fumiya
Kojima, Takeshi
Gambardella, Andrew
Cao, Qi
Iwasawa, Yusuke
Matsuo, Yutaka
Computation and Language
Artificial Intelligence
Recent large language models (LLMs) have demonstrated remarkable generalization abilities in mathematics and logical reasoning tasks. Prior research indicates that LLMs pre-trained with programming language data exhibit high mathematical and reasoning abilities; however, this causal relationship has not been rigorously tested. Our research aims to verify which programming languages and features during pre-training affect logical inference performance. Specifically, we pre-trained decoder-based language models from scratch using datasets from ten programming languages (e.g., Python, C, Java) and three natural language datasets (Wikipedia, Fineweb, C4) under identical conditions. Thereafter, we evaluated the trained models in a few-shot in-context learning setting on logical reasoning tasks: FLD and bAbi, which do not require commonsense or world knowledge. The results demonstrate that nearly all models trained with programming languages consistently outperform those trained with natural languages, indicating that programming languages contain factors that elicit logic inference performance. In addition, we found that models trained with programming languages exhibit a better ability to follow instructions compared to those trained with natural languages. Further analysis reveals that the depth of Abstract Syntax Trees representing parsed results of programs also affects logical reasoning performance. These findings will offer insights into the essential elements of pre-training for acquiring the foundational abilities of LLMs.
title Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.06735