CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Yan, Weixiang, Liu, Haitian, Wang, Yunkun, Li, Yunzhe, Chen, Qian, Wang, Wen, Lin, Tingyu, Zhao, Weishan, Zhu, Li, Sundaram, Hari, Deng, Shuiguang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CodeGlance: Understanding Code Reasoning Challenges in LLMs through Multi-Dimensional Feature Analysis
di: Wang, Yunkun, et al.
Pubblicazione: (2026)
di: Wang, Yunkun, et al.
Pubblicazione: (2026)
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
di: Tian, Yuchen, et al.
Pubblicazione: (2024)
di: Tian, Yuchen, et al.
Pubblicazione: (2024)
Do Code LLMs Understand Design Patterns?
di: Pan, Zhenyu, et al.
Pubblicazione: (2025)
di: Pan, Zhenyu, et al.
Pubblicazione: (2025)
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
di: Gu, Alex, et al.
Pubblicazione: (2024)
di: Gu, Alex, et al.
Pubblicazione: (2024)
GRACE: Graph-Guided Repository-Aware Code Completion through Hierarchical Code Fusion
di: Wang, Xingliang, et al.
Pubblicazione: (2025)
di: Wang, Xingliang, et al.
Pubblicazione: (2025)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
di: Chen, Zaoyu, et al.
Pubblicazione: (2026)
di: Chen, Zaoyu, et al.
Pubblicazione: (2026)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
di: Guo, Chengquan, et al.
Pubblicazione: (2024)
di: Guo, Chengquan, et al.
Pubblicazione: (2024)
SelfPiCo: Self-Guided Partial Code Execution with LLMs
di: Xue, Zhipeng, et al.
Pubblicazione: (2024)
di: Xue, Zhipeng, et al.
Pubblicazione: (2024)
InspectCoder: Dynamic Analysis-Enabled Self Repair through interactive LLM-Debugger Collaboration
di: Wang, Yunkun, et al.
Pubblicazione: (2025)
di: Wang, Yunkun, et al.
Pubblicazione: (2025)
Compiling Code LLMs into Lightweight Executables
di: Shi, Jieke, et al.
Pubblicazione: (2026)
di: Shi, Jieke, et al.
Pubblicazione: (2026)
CodeScore: Evaluating Code Generation by Learning Code Execution
di: Dong, Yihong, et al.
Pubblicazione: (2023)
di: Dong, Yihong, et al.
Pubblicazione: (2023)
Completion by Comprehension: Guiding Code Generation with Multi-Granularity Understanding
di: Zhao, Xinkui, et al.
Pubblicazione: (2025)
di: Zhao, Xinkui, et al.
Pubblicazione: (2025)
ExploraCoder: Advancing code generation for multiple unseen APIs via planning and chained exploration
di: Wang, Yunkun, et al.
Pubblicazione: (2024)
di: Wang, Yunkun, et al.
Pubblicazione: (2024)
CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs
di: Manh, Dung Nguyen, et al.
Pubblicazione: (2024)
di: Manh, Dung Nguyen, et al.
Pubblicazione: (2024)
GrepRAG: An Empirical Study and Optimization of Grep-Like Retrieval for Code Completion
di: Wang, Baoyi, et al.
Pubblicazione: (2026)
di: Wang, Baoyi, et al.
Pubblicazione: (2026)
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
di: Zhao, Yuwei, et al.
Pubblicazione: (2024)
di: Zhao, Yuwei, et al.
Pubblicazione: (2024)
Beyond Retrieval: A Multitask Benchmark and Model for Code Search
di: Xue, Siqiao, et al.
Pubblicazione: (2026)
di: Xue, Siqiao, et al.
Pubblicazione: (2026)
An Execution-Verified Multi-Language Benchmark for Code Semantic Reasoning
di: Li, Yikun, et al.
Pubblicazione: (2026)
di: Li, Yikun, et al.
Pubblicazione: (2026)
Defects4Log: Benchmarking LLMs for Logging Code Defect Detection and Reasoning
di: Wang, Xin, et al.
Pubblicazione: (2025)
di: Wang, Xin, et al.
Pubblicazione: (2025)
Development and Benchmarking of Multilingual Code Clone Detector
di: Zhu, Wenqing, et al.
Pubblicazione: (2024)
di: Zhu, Wenqing, et al.
Pubblicazione: (2024)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
di: Aleithan, Reem, et al.
Pubblicazione: (2024)
di: Aleithan, Reem, et al.
Pubblicazione: (2024)
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
di: Jiang, Xue, et al.
Pubblicazione: (2025)
di: Jiang, Xue, et al.
Pubblicazione: (2025)
Empowering RepoQA-Agent based on Reinforcement Learning Driven by Monte-carlo Tree Search
di: Li, Guochang, et al.
Pubblicazione: (2025)
di: Li, Guochang, et al.
Pubblicazione: (2025)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
di: Xie, Yiqing, et al.
Pubblicazione: (2024)
di: Xie, Yiqing, et al.
Pubblicazione: (2024)
LONGCODEU: Benchmarking Long-Context Language Models on Long Code Understanding
di: Li, Jia, et al.
Pubblicazione: (2025)
di: Li, Jia, et al.
Pubblicazione: (2025)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
di: Zheng, Qinkai, et al.
Pubblicazione: (2023)
di: Zheng, Qinkai, et al.
Pubblicazione: (2023)
Revisiting Evolutionary Program Repair via Code Language Model
di: Wang, Yunan, et al.
Pubblicazione: (2024)
di: Wang, Yunan, et al.
Pubblicazione: (2024)
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
di: Wang, Yubang, et al.
Pubblicazione: (2026)
di: Wang, Yubang, et al.
Pubblicazione: (2026)
Benchmarking LLMs for Fine-Grained Code Review with Enriched Context in Practice
di: Hu, Ruida, et al.
Pubblicazione: (2025)
di: Hu, Ruida, et al.
Pubblicazione: (2025)
COFFE: A Code Efficiency Benchmark for Code Generation
di: Peng, Yun, et al.
Pubblicazione: (2025)
di: Peng, Yun, et al.
Pubblicazione: (2025)
Integrating Symbolic Execution into the Fine-Tuning of Code-Generating LLMs
di: Sakharova, Marina, et al.
Pubblicazione: (2025)
di: Sakharova, Marina, et al.
Pubblicazione: (2025)
Advancing Precise Outline-Conditioned Text Generation with Task Duality and Explicit Outline Control
di: Li, Yunzhe, et al.
Pubblicazione: (2023)
di: Li, Yunzhe, et al.
Pubblicazione: (2023)
LibRec: Benchmarking Retrieval-Augmented LLMs for Library Migration Recommendations
di: Han, Junxiao, et al.
Pubblicazione: (2025)
di: Han, Junxiao, et al.
Pubblicazione: (2025)
Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation
di: Wang, Chong, et al.
Pubblicazione: (2024)
di: Wang, Chong, et al.
Pubblicazione: (2024)
Python Symbolic Execution with LLM-powered Code Generation
di: Wang, Wenhan, et al.
Pubblicazione: (2024)
di: Wang, Wenhan, et al.
Pubblicazione: (2024)
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs
di: Li, Ziyu, et al.
Pubblicazione: (2024)
di: Li, Ziyu, et al.
Pubblicazione: (2024)
Understanding Code Understandability Improvements in Code Reviews
di: Oliveira, Delano, et al.
Pubblicazione: (2024)
di: Oliveira, Delano, et al.
Pubblicazione: (2024)
Assessing Code Understanding in LLMs
di: Laneve, Cosimo, et al.
Pubblicazione: (2025)
di: Laneve, Cosimo, et al.
Pubblicazione: (2025)
Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey
di: Wang, Junqiao, et al.
Pubblicazione: (2024)
di: Wang, Junqiao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
CodeGlance: Understanding Code Reasoning Challenges in LLMs through Multi-Dimensional Feature Analysis
di: Wang, Yunkun, et al.
Pubblicazione: (2026) -
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
di: Tian, Yuchen, et al.
Pubblicazione: (2024) -
Do Code LLMs Understand Design Patterns?
di: Pan, Zhenyu, et al.
Pubblicazione: (2025) -
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
di: Gu, Alex, et al.
Pubblicazione: (2024) -
GRACE: Graph-Guided Repository-Aware Code Completion through Hierarchical Code Fusion
di: Wang, Xingliang, et al.
Pubblicazione: (2025)