CoRe: Benchmarking LLMs Code Reasoning Capabilities through Static Analysis Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Danning, Zheng, Mingwei, Liu, Xuwei, Wang, Jiannan, Wang, Chengpeng, Tan, Lin, Zhang, Xiangyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908771776724992
author Xie, Danning
Zheng, Mingwei
Liu, Xuwei
Wang, Jiannan
Wang, Chengpeng
Tan, Lin
Zhang, Xiangyu
author_facet Xie, Danning
Zheng, Mingwei
Liu, Xuwei
Wang, Jiannan
Wang, Chengpeng
Tan, Lin
Zhang, Xiangyu
contents Large language models (LLMs) have been widely adopted across diverse domains of software engineering, such as code generation, program repair, and vulnerability detection. These applications require understanding beyond surface-level code patterns: value propagation, control flow, and interdependence between program elements. However, existing benchmarks primarily evaluate end-to-end outcomes, such as whether code is correctly repaired or generated, leaving the models' ability for program semantic reasoning underexplored. This work presents CORE, a high-quality, human-verified benchmark designed to evaluate LLMs on fundamental static analysis tasks. CORE includes 12,553 task instances spanning data dependency, control dependency, and information flow across programs written in C/C++, Java, and Python. To ensure semantic diversity and reasoning complexity, we propose a semantics-aware diverse sampling strategy that selects targets and task instances based on structural coverage and dependency depth. We evaluate 10 mainstream LLMs and show that, while they perform well at identifying dependencies, models still struggle with tasks that require deeper semantic understanding and multi-step reasoning. We further conduct qualitative analyses to uncover key challenges, such as complex control structures and backward dependency patterns, offering insights into improving LLMs' code reasoning capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05269
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoRe: Benchmarking LLMs Code Reasoning Capabilities through Static Analysis Tasks
Xie, Danning
Zheng, Mingwei
Liu, Xuwei
Wang, Jiannan
Wang, Chengpeng
Tan, Lin
Zhang, Xiangyu
Software Engineering
Artificial Intelligence
Large language models (LLMs) have been widely adopted across diverse domains of software engineering, such as code generation, program repair, and vulnerability detection. These applications require understanding beyond surface-level code patterns: value propagation, control flow, and interdependence between program elements. However, existing benchmarks primarily evaluate end-to-end outcomes, such as whether code is correctly repaired or generated, leaving the models' ability for program semantic reasoning underexplored. This work presents CORE, a high-quality, human-verified benchmark designed to evaluate LLMs on fundamental static analysis tasks. CORE includes 12,553 task instances spanning data dependency, control dependency, and information flow across programs written in C/C++, Java, and Python. To ensure semantic diversity and reasoning complexity, we propose a semantics-aware diverse sampling strategy that selects targets and task instances based on structural coverage and dependency depth. We evaluate 10 mainstream LLMs and show that, while they perform well at identifying dependencies, models still struggle with tasks that require deeper semantic understanding and multi-step reasoning. We further conduct qualitative analyses to uncover key challenges, such as complex control structures and backward dependency patterns, offering insights into improving LLMs' code reasoning capabilities.
title CoRe: Benchmarking LLMs Code Reasoning Capabilities through Static Analysis Tasks
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2507.05269