CoRe: Benchmarking LLMs Code Reasoning Capabilities through Static Analysis Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Danning, Zheng, Mingwei, Liu, Xuwei, Wang, Jiannan, Wang, Chengpeng, Tan, Lin, Zhang, Xiangyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Validating Network Protocol Parsers with Traceable RFC Document Interpretation
by: Zheng, Mingwei, et al.
Published: (2025)
by: Zheng, Mingwei, et al.
Published: (2025)
NESA: Relational Neuro-Symbolic Static Program Analysis
by: Wang, Chengpeng, et al.
Published: (2024)
by: Wang, Chengpeng, et al.
Published: (2024)
Raw Pointer Rewriting with LLMs for Translating C to Safer Rust
by: Gao, Yifei, et al.
Published: (2025)
by: Gao, Yifei, et al.
Published: (2025)
Large Language Models for Validating Network Protocol Parsers
by: Zheng, Mingwei, et al.
Published: (2025)
by: Zheng, Mingwei, et al.
Published: (2025)
RFCAudit: An LLM Agent for Functional Bug Detection in Network Protocols
by: Zheng, Mingwei, et al.
Published: (2025)
by: Zheng, Mingwei, et al.
Published: (2025)
CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs
by: Manh, Dung Nguyen, et al.
Published: (2024)
by: Manh, Dung Nguyen, et al.
Published: (2024)
STALL+: Boosting LLM-based Repository-level Code Completion with Static Analysis
by: Liu, Junwei, et al.
Published: (2024)
by: Liu, Junwei, et al.
Published: (2024)
Exploring the Capabilities of LLMs for Code Change Related Tasks
by: Fan, Lishui, et al.
Published: (2024)
by: Fan, Lishui, et al.
Published: (2024)
Do Code LLMs Do Static Analysis?
by: Su, Chia-Yi, et al.
Published: (2025)
by: Su, Chia-Yi, et al.
Published: (2025)
CodeGlance: Understanding Code Reasoning Challenges in LLMs through Multi-Dimensional Feature Analysis
by: Wang, Yunkun, et al.
Published: (2026)
by: Wang, Yunkun, et al.
Published: (2026)
How Effective are Large Language Models in Generating Software Specifications?
by: Xie, Danning, et al.
Published: (2023)
by: Xie, Danning, et al.
Published: (2023)
EffiReasonTrans: RL-Optimized Reasoning for Code Translation
by: Wang, Yanlin, et al.
Published: (2025)
by: Wang, Yanlin, et al.
Published: (2025)
A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback
by: Duan, Guoliang, et al.
Published: (2025)
by: Duan, Guoliang, et al.
Published: (2025)
RustRepoTrans: Repository-level Code Translation Benchmark Targeting Rust
by: Ou, Guangsheng, et al.
Published: (2024)
by: Ou, Guangsheng, et al.
Published: (2024)
Defects4Log: Benchmarking LLMs for Logging Code Defect Detection and Reasoning
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Identifying Adversary Tactics and Techniques in Malware Binaries with an LLM Agent
by: Xuan, Zhou, et al.
Published: (2026)
by: Xuan, Zhou, et al.
Published: (2026)
LLMs are Bug Replicators: An Empirical Study on LLMs' Capability in Completing Bug-prone Code
by: Guo, Liwei, et al.
Published: (2025)
by: Guo, Liwei, et al.
Published: (2025)
Static Code Analysis with CodeChecker
by: Horvath, Gabor, et al.
Published: (2024)
by: Horvath, Gabor, et al.
Published: (2024)
Evolving Triple Knowledge-Augmented LLMs for Code Translation in Repository Context
by: Ou, Guangsheng, et al.
Published: (2025)
by: Ou, Guangsheng, et al.
Published: (2025)
HintPilot: LLM-based Compiler Hint Synthesis for Code Optimization
by: Jiang, Hanyun, et al.
Published: (2026)
by: Jiang, Hanyun, et al.
Published: (2026)
Integrating Static Code Analysis Toolchains
by: Kern, Matthias, et al.
Published: (2024)
by: Kern, Matthias, et al.
Published: (2024)
Position: Intelligent Coding Systems Should Write Programs with Justifications
by: Xu, Xiangzhe, et al.
Published: (2025)
by: Xu, Xiangzhe, et al.
Published: (2025)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
by: Aleithan, Reem, et al.
Published: (2024)
by: Aleithan, Reem, et al.
Published: (2024)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
An Empirical Study of False Negatives and Positives of Static Code Analyzers From the Perspective of Historical Issues
by: Cui, Han, et al.
Published: (2024)
by: Cui, Han, et al.
Published: (2024)
QLPro: Automated Code Vulnerability Discovery via LLM and Static Code Analysis Integration
by: Hu, Junze, et al.
Published: (2025)
by: Hu, Junze, et al.
Published: (2025)
Large Language Models Versus Static Code Analysis Tools: A Systematic Benchmark for Vulnerability Detection
by: Gnieciak, Damian, et al.
Published: (2025)
by: Gnieciak, Damian, et al.
Published: (2025)
UniCode: Augmenting Evaluation for Code Reasoning
by: Zheng, Xinyue, et al.
Published: (2025)
by: Zheng, Xinyue, et al.
Published: (2025)
DocTer: Documentation Guided Fuzzing for Testing Deep Learning API Functions
by: Xie, Danning, et al.
Published: (2021)
by: Xie, Danning, et al.
Published: (2021)
The Emergence of Large Language Models in Static Analysis: A First Look through Micro-Benchmarks
by: Venkatesh, Ashwin Prasad Shivarpatna, et al.
Published: (2024)
by: Venkatesh, Ashwin Prasad Shivarpatna, et al.
Published: (2024)
BugScope: Learn to Find Bugs Like Human
by: Guo, Jinyao, et al.
Published: (2025)
by: Guo, Jinyao, et al.
Published: (2025)
FeedbackEval: A Benchmark for Evaluating Large Language Models in Feedback-Driven Code Repair Tasks
by: Dai, Dekun, et al.
Published: (2025)
by: Dai, Dekun, et al.
Published: (2025)
CoRE: A Fine-Grained Code Reasoning Benchmark Beyond Output Prediction
by: Gao, Jun, et al.
Published: (2026)
by: Gao, Jun, et al.
Published: (2026)
RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation
by: Liang, Linxi, et al.
Published: (2025)
by: Liang, Linxi, et al.
Published: (2025)
ZeroFalse: Improving Precision in Static Analysis with LLMs
by: Iranmanesh, Mohsen, et al.
Published: (2025)
by: Iranmanesh, Mohsen, et al.
Published: (2025)
A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
by: Liu, Shuhan, et al.
Published: (2026)
by: Liu, Shuhan, et al.
Published: (2026)
Yet Even Less Is Even Better For Agentic, Reasoning, and Coding LLMs
by: CodeArts Model Team, et al.
Published: (2026)
by: CodeArts Model Team, et al.
Published: (2026)
What Makes Good In-context Demonstrations for Code Intelligence Tasks with LLMs?
by: Gao, Shuzheng, et al.
Published: (2023)
by: Gao, Shuzheng, et al.
Published: (2023)
Symbol Preference Aware Generative Models for Recovering Variable Names from Stripped Binary
by: Xu, Xiangzhe, et al.
Published: (2023)
by: Xu, Xiangzhe, et al.
Published: (2023)
Unmasking the Genuine Type Inference Capabilities of LLMs for Java Code Snippets
by: Dong, Yiwen, et al.
Published: (2025)
by: Dong, Yiwen, et al.
Published: (2025)
Similar Items
-
Validating Network Protocol Parsers with Traceable RFC Document Interpretation
by: Zheng, Mingwei, et al.
Published: (2025) -
NESA: Relational Neuro-Symbolic Static Program Analysis
by: Wang, Chengpeng, et al.
Published: (2024) -
Raw Pointer Rewriting with LLMs for Translating C to Safer Rust
by: Gao, Yifei, et al.
Published: (2025) -
Large Language Models for Validating Network Protocol Parsers
by: Zheng, Mingwei, et al.
Published: (2025) -
RFCAudit: An LLM Agent for Functional Bug Detection in Network Protocols
by: Zheng, Mingwei, et al.
Published: (2025)