Strengthening Programming Comprehension in Large Language Models through Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ren, Xiaoning, Hu, Qiang, Ma, Wei, Li, Yan, Zhang, Yao, Jiang, Lingxiao, Xue, Yinxing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916905112043520
author Ren, Xiaoning
Hu, Qiang
Ma, Wei
Li, Yan
Zhang, Yao
Jiang, Lingxiao
Xue, Yinxing
author_facet Ren, Xiaoning
Hu, Qiang
Ma, Wei
Li, Yan
Zhang, Yao
Jiang, Lingxiao
Xue, Yinxing
contents Large language models (LLMs) have recently shown impressive results on diverse code-related tasks, benefiting from large-scale training and instruction tuning. However, studies reveal that their grasp of fundamental programming concepts, such as data flow and control flow, remains shallow, leading to fragile performance when code requires deeper reasoning. This limitation restricts the practical adoption of LLMs in real-world software development. To address this issue, this work introduces a counterfactual code augmentation framework combined with concept-aware tuning, designed to guide LLMs toward stronger conceptual understanding. Comprehensive evaluation across multiple models and benchmarks demonstrates the effectiveness of the proposed approach.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12620
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Strengthening Programming Comprehension in Large Language Models through Code Generation
Ren, Xiaoning
Hu, Qiang
Ma, Wei
Li, Yan
Zhang, Yao
Jiang, Lingxiao
Xue, Yinxing
Software Engineering
Programming Languages
Large language models (LLMs) have recently shown impressive results on diverse code-related tasks, benefiting from large-scale training and instruction tuning. However, studies reveal that their grasp of fundamental programming concepts, such as data flow and control flow, remains shallow, leading to fragile performance when code requires deeper reasoning. This limitation restricts the practical adoption of LLMs in real-world software development. To address this issue, this work introduces a counterfactual code augmentation framework combined with concept-aware tuning, designed to guide LLMs toward stronger conceptual understanding. Comprehensive evaluation across multiple models and benchmarks demonstrates the effectiveness of the proposed approach.
title Strengthening Programming Comprehension in Large Language Models through Code Generation
topic Software Engineering
Programming Languages
url https://arxiv.org/abs/2508.12620