Rigor, Reliability, and Reproducibility Matter: A Decade-Scale Survey of 572 Code Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Jialun, Chan, Yuk-Kit, Ling, Zixuan, Wang, Wenxuan, Li, Shuqing, Liu, Mingwei, Qiao, Ruixi, Han, Yuting, Wang, Chaozheng, Yu, Boxi, He, Pinjia, Wang, Shuai, Zheng, Zibin, Lyu, Michael R., Cheung, Shing-Chi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
by: Yu, Boxi, et al.
Published: (2025)
by: Yu, Boxi, et al.
Published: (2025)
Learning to Ask: When LLM Agents Meet Unclear Instruction
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Automatically Generating UI Code from Screenshot: A Divide-and-Conquer-Based Approach
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
by: Yu, Boxi, et al.
Published: (2026)
by: Yu, Boxi, et al.
Published: (2026)
Concerned with Data Contamination? Assessing Countermeasures in Code Language Model
by: Cao, Jialun, et al.
Published: (2024)
by: Cao, Jialun, et al.
Published: (2024)
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
by: Wang, Liwen, et al.
Published: (2025)
by: Wang, Liwen, et al.
Published: (2025)
Understanding and Bridging the Planner-Coder Gap: A Systematic Study on the Robustness of Multi-Agent Systems for Code Generation
by: Lyu, Zongyi, et al.
Published: (2025)
by: Lyu, Zongyi, et al.
Published: (2025)
Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
ReuseDroid: A VLM-empowered Android UI Test Migrator Boosted by Active Feedback
by: Li, Xiaolei, et al.
Published: (2025)
by: Li, Xiaolei, et al.
Published: (2025)
What Builds Effective In-Context Examples for Code Generation?
by: Li, Dongze, et al.
Published: (2025)
by: Li, Dongze, et al.
Published: (2025)
ModelWisdom: An Integrated Toolkit for TLA+ Model Visualization, Digest and Repair
by: Chen, Zhiyong, et al.
Published: (2026)
by: Chen, Zhiyong, et al.
Published: (2026)
An Empirical Study on Package-Level Deprecation in Python Ecosystem
by: Zhong, Zhiqing, et al.
Published: (2024)
by: Zhong, Zhiqing, et al.
Published: (2024)
Exploring Multi-Lingual Bias of Large Code Models in Code Generation
by: Wang, Chaozheng, et al.
Published: (2024)
by: Wang, Chaozheng, et al.
Published: (2024)
Metamorphic Testing for Audio Content Moderation Software
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
3D Software Synthesis Guided by Constraint-Expressive Intermediate Representation
by: Li, Shuqing, et al.
Published: (2025)
by: Li, Shuqing, et al.
Published: (2025)
CAM: A Causality-based Analysis Framework for Multi-Agent Code Generation Systems
by: Lyu, Zongyi, et al.
Published: (2026)
by: Lyu, Zongyi, et al.
Published: (2026)
RC-positivity, comparison theorems and prescribed Hermitian-Yang-Mills tensors I
by: Wang, Mingwei, et al.
Published: (2026)
by: Wang, Mingwei, et al.
Published: (2026)
When LLMs Meet API Documentation: Can Retrieval Augmentation Aid Code Generation Just as It Helps Developers?
by: Chen, Jingyi, et al.
Published: (2025)
by: Chen, Jingyi, et al.
Published: (2025)
Can Large Language Models Model Programs Formally?
by: Chen, Zhiyong, et al.
Published: (2026)
by: Chen, Zhiyong, et al.
Published: (2026)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
by: Cao, Jialun, et al.
Published: (2024)
by: Cao, Jialun, et al.
Published: (2024)
Grounded GUI Understanding for Vision-Based Spatial Intelligent Agent: Exemplified by Extended Reality Apps
by: Li, Shuqing, et al.
Published: (2024)
by: Li, Shuqing, et al.
Published: (2024)
A Systematic Evaluation of Large Code Models in API Suggestion: When, Which, and How
by: Wang, Chaozheng, et al.
Published: (2024)
by: Wang, Chaozheng, et al.
Published: (2024)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
A Recursive Algorithm for Multi-Coefficient Inversion in Nonlinear Helmholtz Equations
by: Lu, Shuai, et al.
Published: (2025)
by: Lu, Shuai, et al.
Published: (2025)
WARBENCH: A Comprehensive Benchmark for Evaluating LLMs in Military Decision-Making
by: Li, Zongjie, et al.
Published: (2026)
by: Li, Zongjie, et al.
Published: (2026)
CODECLEANER: Elevating Standards with A Robust Data Contamination Mitigation Toolkit
by: Cao, Jialun, et al.
Published: (2024)
by: Cao, Jialun, et al.
Published: (2024)
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
by: Wu, Jiarong, et al.
Published: (2025)
by: Wu, Jiarong, et al.
Published: (2025)
SEER: Enhancing Chain-of-Thought Code Generation through Self-Exploring Deep Reasoning
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
Big data approach in the field of gastric and colorectal cancer research
by: Ka Shing Cheung
Published: (2024)
by: Ka Shing Cheung
Published: (2024)
The Use of Digital Watermarking for Intelligence Multimedia Document Distribution
by: Shing-Chi Cheung
Published: (2008)
by: Shing-Chi Cheung
Published: (2008)
RepoDoc: A Knowledge Graph-Based Framework to Automatic Documentation Generation and Incremental Updates
by: Xu, Dong, et al.
Published: (2026)
by: Xu, Dong, et al.
Published: (2026)
Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition
by: Liu, Mingwei, et al.
Published: (2026)
by: Liu, Mingwei, et al.
Published: (2026)
Dynamic analysis enhances issue resolution
by: Liu, Mingwei, et al.
Published: (2026)
by: Liu, Mingwei, et al.
Published: (2026)
A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback
by: Duan, Guoliang, et al.
Published: (2025)
by: Duan, Guoliang, et al.
Published: (2025)
Position Paper: Programming Language Techniques for Bridging LLM Code Generation Semantic Gaps
by: Du, Yalong, et al.
Published: (2025)
by: Du, Yalong, et al.
Published: (2025)
Can Emulating Semantic Translation Help LLMs with Code Translation? A Study Based on Pseudocode
by: Chen, Songqiang, et al.
Published: (2025)
by: Chen, Songqiang, et al.
Published: (2025)
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
by: Li, Meiziniu, et al.
Published: (2024)
by: Li, Meiziniu, et al.
Published: (2024)
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
by: Li, Meiziniu, et al.
Published: (2022)
by: Li, Meiziniu, et al.
Published: (2022)
RC-positivity, comparison theorems and prescribed Hermitian-Yang-Mills tensors II
by: Fan, Jiaxuan, et al.
Published: (2026)
by: Fan, Jiaxuan, et al.
Published: (2026)
SAMUeL: Efficient Vocal-Conditioned Music Generation via Soft Alignment Attention and Latent Diffusion
by: Cheung, Hei Shing, et al.
Published: (2025)
by: Cheung, Hei Shing, et al.
Published: (2025)
Similar Items
-
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
by: Yu, Boxi, et al.
Published: (2025) -
Learning to Ask: When LLM Agents Meet Unclear Instruction
by: Wang, Wenxuan, et al.
Published: (2024) -
Automatically Generating UI Code from Screenshot: A Divide-and-Conquer-Based Approach
by: Wan, Yuxuan, et al.
Published: (2024) -
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
by: Yu, Boxi, et al.
Published: (2026) -
Concerned with Data Contamination? Assessing Countermeasures in Code Language Model
by: Cao, Jialun, et al.
Published: (2024)