AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
Fuente:
arXiv
Salvato in:
| Autori principali: | Chou, Jason, Liu, Ao, Deng, Yuchi, Zeng, Zhiying, Zhang, Tao, Zhu, Haotian, Cai, Jianwei, Mao, Yue, Zhang, Chenchen, Tan, Lingyun, Xu, Ziyan, Zhai, Bohui, Liu, Hengyi, Zhu, Speed, Zhou, Wiggin, Lian, Fengzong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generation
di: Zhu, Speed, et al.
Pubblicazione: (2025)
di: Zhu, Speed, et al.
Pubblicazione: (2025)
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
di: Zhang, Chenchen, et al.
Pubblicazione: (2025)
di: Zhang, Chenchen, et al.
Pubblicazione: (2025)
ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding
di: Li, Yuhang, et al.
Pubblicazione: (2025)
di: Li, Yuhang, et al.
Pubblicazione: (2025)
AutoICE: Automatically Synthesizing Verifiable C Code via LLM-driven Evolution
di: Luo, Weilin, et al.
Pubblicazione: (2025)
di: Luo, Weilin, et al.
Pubblicazione: (2025)
RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations
di: Li, Hanyu, et al.
Pubblicazione: (2026)
di: Li, Hanyu, et al.
Pubblicazione: (2026)
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
di: Zheng, Dewu, et al.
Pubblicazione: (2024)
di: Zheng, Dewu, et al.
Pubblicazione: (2024)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
di: Huang, Dong, et al.
Pubblicazione: (2024)
di: Huang, Dong, et al.
Pubblicazione: (2024)
The Gap Between Principle and Practice of Lossy Image Coding
di: Zhang, Haotian, et al.
Pubblicazione: (2025)
di: Zhang, Haotian, et al.
Pubblicazione: (2025)
Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
di: Liu, Fang, et al.
Pubblicazione: (2024)
di: Liu, Fang, et al.
Pubblicazione: (2024)
Codes for Limited-Magnitude Probability Error in DNA Storage
di: Zhang, Wenkai, et al.
Pubblicazione: (2024)
di: Zhang, Wenkai, et al.
Pubblicazione: (2024)
Novel Translations
di: Wiggin, Bethany
Pubblicazione: (2023)
di: Wiggin, Bethany
Pubblicazione: (2023)
Enduring colonial legacies in Philadelphia
di: Bethany Wiggin
Pubblicazione: (2024)
di: Bethany Wiggin
Pubblicazione: (2024)
Learning to Guarantee Type Correctness in Code Generation through Type-Guided Program Synthesis
di: Huang, Zhechong, et al.
Pubblicazione: (2025)
di: Huang, Zhechong, et al.
Pubblicazione: (2025)
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
SecureReviewer: Enhancing Large Language Models for Secure Code Review through Secure-aware Fine-tuning
di: Liu, Fang, et al.
Pubblicazione: (2025)
di: Liu, Fang, et al.
Pubblicazione: (2025)
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
di: Zhu, Xiao, et al.
Pubblicazione: (2026)
di: Zhu, Xiao, et al.
Pubblicazione: (2026)
AACR-Bench: Evaluating Automatic Code Review with Holistic Repository-Level Context
di: Zhang, Lei, et al.
Pubblicazione: (2026)
di: Zhang, Lei, et al.
Pubblicazione: (2026)
CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models
di: Zhang, Alexander, et al.
Pubblicazione: (2025)
di: Zhang, Alexander, et al.
Pubblicazione: (2025)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
di: Wang, Yanli, et al.
Pubblicazione: (2024)
di: Wang, Yanli, et al.
Pubblicazione: (2024)
A Unified Error Correction Code for Universal Quantum Computing with Identical Particles
di: Wu, S. L., et al.
Pubblicazione: (2026)
di: Wu, S. L., et al.
Pubblicazione: (2026)
Multilingual Multimodal Software Developer for Code Generation
di: Chai, Linzheng, et al.
Pubblicazione: (2025)
di: Chai, Linzheng, et al.
Pubblicazione: (2025)
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games
di: Zhang, Wenyu, et al.
Pubblicazione: (2026)
di: Zhang, Wenyu, et al.
Pubblicazione: (2026)
Distributed Observer Design over Directed Switching Topologies
di: Xu, Haotian, et al.
Pubblicazione: (2024)
di: Xu, Haotian, et al.
Pubblicazione: (2024)
New Hampshire: The Automated Information System.
di: Wiggin, Kendall F.
Pubblicazione: (1996)
di: Wiggin, Kendall F.
Pubblicazione: (1996)
Gateway 2000: A Strategic Plan for the New Hampshire Automated Information System.
di: Wiggin, Kendall F.
Pubblicazione: (1993)
di: Wiggin, Kendall F.
Pubblicazione: (1993)
SWE Context Bench: A Benchmark for Context Learning in Coding
di: Zhu, Jiayuan, et al.
Pubblicazione: (2026)
di: Zhu, Jiayuan, et al.
Pubblicazione: (2026)
Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
di: Zhang, Zhe, et al.
Pubblicazione: (2025)
di: Zhang, Zhe, et al.
Pubblicazione: (2025)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
di: Yu, Boxi, et al.
Pubblicazione: (2025)
di: Yu, Boxi, et al.
Pubblicazione: (2025)
Extended p-median problems for balancing service efficiency and equality
di: Kong, Yunfeng, et al.
Pubblicazione: (2023)
di: Kong, Yunfeng, et al.
Pubblicazione: (2023)
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
di: Zhang, Shudan, et al.
Pubblicazione: (2024)
di: Zhang, Shudan, et al.
Pubblicazione: (2024)
DSCodeBench: A Realistic Benchmark for Data Science Code Generation
di: Ouyang, Shuyin, et al.
Pubblicazione: (2025)
di: Ouyang, Shuyin, et al.
Pubblicazione: (2025)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
di: Wang, Peiding, et al.
Pubblicazione: (2025)
di: Wang, Peiding, et al.
Pubblicazione: (2025)
CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
di: Guo, Jiawei, et al.
Pubblicazione: (2024)
di: Guo, Jiawei, et al.
Pubblicazione: (2024)
AutoMCQ -- Automatically Generate Code Comprehension Questions using GenAI
di: Goodfellow, Martin, et al.
Pubblicazione: (2025)
di: Goodfellow, Martin, et al.
Pubblicazione: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
di: Li, Jia, et al.
Pubblicazione: (2024)
di: Li, Jia, et al.
Pubblicazione: (2024)
Column Twisted Reed-Solomon Codes as MDS Codes
di: Liu, Wei, et al.
Pubblicazione: (2025)
di: Liu, Wei, et al.
Pubblicazione: (2025)
AutoBench: Automatic Testbench Generation and Evaluation Using LLMs for HDL Design
di: Qiu, Ruidi, et al.
Pubblicazione: (2024)
di: Qiu, Ruidi, et al.
Pubblicazione: (2024)
CodeMEM: AST-Guided Adaptive Memory for Repository-Level Iterative Code Generation
di: Wang, Peiding, et al.
Pubblicazione: (2026)
di: Wang, Peiding, et al.
Pubblicazione: (2026)
Automatically Recommend Code Updates: Are We There Yet?
di: Liu, Yue, et al.
Pubblicazione: (2022)
di: Liu, Yue, et al.
Pubblicazione: (2022)
Documenti analoghi
-
DRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generation
di: Zhu, Speed, et al.
Pubblicazione: (2025) -
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
di: Zhang, Chenchen, et al.
Pubblicazione: (2025) -
ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding
di: Li, Yuhang, et al.
Pubblicazione: (2025) -
AutoICE: Automatically Synthesizing Verifiable C Code via LLM-driven Evolution
di: Luo, Weilin, et al.
Pubblicazione: (2025) -
RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations
di: Li, Hanyu, et al.
Pubblicazione: (2026)