Evaluating the Generalization Capabilities of Large Language Models on Code Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866910905718013952 |
|---|---|
| author | Yang, Rem Dai, Julian Vasilakis, Nikos Rinard, Martin |
| author_facet | Yang, Rem Dai, Julian Vasilakis, Nikos Rinard, Martin |
| contents | We assess how the code reasoning abilities of large language models (LLMs) generalize to different kinds of programs. We present techniques for obtaining in- and out-of-distribution programs with different characteristics: code sampled from a domain-specific language, code automatically generated by an LLM, code collected from competitive programming contests, and mutated versions of these programs. We also present an experimental methodology for evaluating LLM generalization by comparing their performance on these programs. We perform an extensive evaluation across 10 state-of-the-art models from the past year, obtaining insights into their generalization capabilities over time and across different classes of programs. Our results highlight that while earlier models exhibit behavior consistent with pattern matching, the latest models exhibit strong generalization abilities on code reasoning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_05518 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Evaluating the Generalization Capabilities of Large Language Models on Code Reasoning Yang, Rem Dai, Julian Vasilakis, Nikos Rinard, Martin Software Engineering Computation and Language Machine Learning We assess how the code reasoning abilities of large language models (LLMs) generalize to different kinds of programs. We present techniques for obtaining in- and out-of-distribution programs with different characteristics: code sampled from a domain-specific language, code automatically generated by an LLM, code collected from competitive programming contests, and mutated versions of these programs. We also present an experimental methodology for evaluating LLM generalization by comparing their performance on these programs. We perform an extensive evaluation across 10 state-of-the-art models from the past year, obtaining insights into their generalization capabilities over time and across different classes of programs. Our results highlight that while earlier models exhibit behavior consistent with pattern matching, the latest models exhibit strong generalization abilities on code reasoning. |
| title | Evaluating the Generalization Capabilities of Large Language Models on Code Reasoning |
| topic | Software Engineering Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2504.05518 |