ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex Code
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Jia, Liu, Jiachen, Gao, Cuiyun, Chong, Chun Yong, Wang, Chaozheng, Gao, Shan, Xia, Xin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Systematic Evaluation of Large Code Models in API Suggestion: When, Which, and How
von: Wang, Chaozheng, et al.
Veröffentlicht: (2024)
von: Wang, Chaozheng, et al.
Veröffentlicht: (2024)
SEER: Enhancing Chain-of-Thought Code Generation through Self-Exploring Deep Reasoning
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
SR-Eval: Evaluating LLMs on Code Generation under Stepwise Requirement Refinement
von: Zhan, Zexun, et al.
Veröffentlicht: (2025)
von: Zhan, Zexun, et al.
Veröffentlicht: (2025)
AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
Exploring Multi-Lingual Bias of Large Code Models in Code Generation
von: Wang, Chaozheng, et al.
Veröffentlicht: (2024)
von: Wang, Chaozheng, et al.
Veröffentlicht: (2024)
Cascaded Code Editing: Large-Small Model Collaboration for Effective and Efficient Code Editing
von: Wang, Chaozheng, et al.
Veröffentlicht: (2026)
von: Wang, Chaozheng, et al.
Veröffentlicht: (2026)
FeedbackEval: A Benchmark for Evaluating Large Language Models in Feedback-Driven Code Repair Tasks
von: Dai, Dekun, et al.
Veröffentlicht: (2025)
von: Dai, Dekun, et al.
Veröffentlicht: (2025)
A Deep Dive into Retrieval-Augmented Generation for Code Completion: Experience on WeChat
von: Yang, Zezhou, et al.
Veröffentlicht: (2025)
von: Yang, Zezhou, et al.
Veröffentlicht: (2025)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
von: Wu, Fan, et al.
Veröffentlicht: (2026)
von: Wu, Fan, et al.
Veröffentlicht: (2026)
CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation
von: Wang, Xinchen, et al.
Veröffentlicht: (2025)
von: Wang, Xinchen, et al.
Veröffentlicht: (2025)
Automated Prompt Generation for Code Intelligence: An Empirical study and Experience in WeChat
von: Ji, Kexing, et al.
Veröffentlicht: (2025)
von: Ji, Kexing, et al.
Veröffentlicht: (2025)
RAG or Fine-tuning? A Comparative Study on LCMs-based Code Completion in Industry
von: Wang, Chaozheng, et al.
Veröffentlicht: (2025)
von: Wang, Chaozheng, et al.
Veröffentlicht: (2025)
RepoMasterEval: Evaluating Code Completion via Real-World Repositories
von: Wu, Qinyun, et al.
Veröffentlicht: (2024)
von: Wu, Qinyun, et al.
Veröffentlicht: (2024)
An Empirical Study of Knowledge Distillation for Code Understanding Tasks
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
A Roadmap on Modern Code Review: Challenges and Opportunities
von: Yang, Zezhou, et al.
Veröffentlicht: (2024)
von: Yang, Zezhou, et al.
Veröffentlicht: (2024)
Code Digital Twin: Empowering LLMs with Tacit Knowledge for Complex Software Development
von: Peng, Xin, et al.
Veröffentlicht: (2025)
von: Peng, Xin, et al.
Veröffentlicht: (2025)
The Prompt Alchemist: Automated LLM-Tailored Prompt Optimization for Test Case Generation
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
An Empirical Study of Retrieval-Augmented Code Generation: Challenges and Opportunities
von: Yang, Zezhou, et al.
Veröffentlicht: (2025)
von: Yang, Zezhou, et al.
Veröffentlicht: (2025)
Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code Models
von: Gao, Shuzheng, et al.
Veröffentlicht: (2024)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2024)
Code Digital Twin: A Knowledge Infrastructure for AI-Assisted Complex Software Development
von: Peng, Xin, et al.
Veröffentlicht: (2025)
von: Peng, Xin, et al.
Veröffentlicht: (2025)
The Current Challenges of Software Engineering in the Era of Large Language Models
von: Gao, Cuiyun, et al.
Veröffentlicht: (2024)
von: Gao, Cuiyun, et al.
Veröffentlicht: (2024)
SolEval: Benchmarking Large Language Models for Repository-level Solidity Code Generation
von: Peng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Peng, Zhiyuan, et al.
Veröffentlicht: (2025)
Dependency-Guided Repository-Level C-to-Rust Translation with Reinforcement Alignment
von: Feng, Jia, et al.
Veröffentlicht: (2026)
von: Feng, Jia, et al.
Veröffentlicht: (2026)
A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
von: Gao, Cuiyun, et al.
Veröffentlicht: (2025)
von: Gao, Cuiyun, et al.
Veröffentlicht: (2025)
Uncovering Weaknesses in Neural Code Generation
von: Lian, Xiaoli, et al.
Veröffentlicht: (2024)
von: Lian, Xiaoli, et al.
Veröffentlicht: (2024)
What Makes Good In-context Demonstrations for Code Intelligence Tasks with LLMs?
von: Gao, Shuzheng, et al.
Veröffentlicht: (2023)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2023)
Benchmarking LLMs for Fine-Grained Code Review with Enriched Context in Practice
von: Hu, Ruida, et al.
Veröffentlicht: (2025)
von: Hu, Ruida, et al.
Veröffentlicht: (2025)
ClassEval-T: Evaluating Large Language Models in Class-Level Code Translation
von: Xue, Pengyu, et al.
Veröffentlicht: (2024)
von: Xue, Pengyu, et al.
Veröffentlicht: (2024)
SpecEval: Evaluating Code Comprehension in Large Language Models via Program Specifications
von: Ma, Lezhi, et al.
Veröffentlicht: (2024)
von: Ma, Lezhi, et al.
Veröffentlicht: (2024)
CodeRepoQA: A Large-scale Benchmark for Software Engineering Question Answering
von: Hu, Ruida, et al.
Veröffentlicht: (2024)
von: Hu, Ruida, et al.
Veröffentlicht: (2024)
AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
von: Zhang, Tanghaoran, et al.
Veröffentlicht: (2026)
von: Zhang, Tanghaoran, et al.
Veröffentlicht: (2026)
Empirical Study of Code Large Language Models for Binary Security Patch Detection
von: Li, Qingyuan, et al.
Veröffentlicht: (2025)
von: Li, Qingyuan, et al.
Veröffentlicht: (2025)
Search-Based LLMs for Code Optimization
von: Gao, Shuzheng, et al.
Veröffentlicht: (2024)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2024)
Bridge and Hint: Extending Pre-trained Language Models for Long-Range Code
von: Chen, Yujia, et al.
Veröffentlicht: (2024)
von: Chen, Yujia, et al.
Veröffentlicht: (2024)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS
von: Xie, Bang, et al.
Veröffentlicht: (2026)
von: Xie, Bang, et al.
Veröffentlicht: (2026)
Clean Code, Better Models: Enhancing LLM Performance with Smell-Cleaned Dataset
von: Xue, Zhipeng, et al.
Veröffentlicht: (2025)
von: Xue, Zhipeng, et al.
Veröffentlicht: (2025)
A Survey on Evaluating Large Language Models in Code Generation Tasks
von: Chen, Liguo, et al.
Veröffentlicht: (2024)
von: Chen, Liguo, et al.
Veröffentlicht: (2024)
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
von: Liu, Kaiyuan, et al.
Veröffentlicht: (2025)
von: Liu, Kaiyuan, et al.
Veröffentlicht: (2025)
Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval
von: Wang, Jiexin, et al.
Veröffentlicht: (2024)
von: Wang, Jiexin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Systematic Evaluation of Large Code Models in API Suggestion: When, Which, and How
von: Wang, Chaozheng, et al.
Veröffentlicht: (2024) -
SEER: Enhancing Chain-of-Thought Code Generation through Self-Exploring Deep Reasoning
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025) -
SR-Eval: Evaluating LLMs on Code Generation under Stepwise Requirement Refinement
von: Zhan, Zexun, et al.
Veröffentlicht: (2025) -
AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025) -
Exploring Multi-Lingual Bias of Large Code Models in Code Generation
von: Wang, Chaozheng, et al.
Veröffentlicht: (2024)