Salvato in:
| Autori principali: | Szych, Joanna, Schwerk, Anne |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2605.09059 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Benchmarking and Studying the LLM-based Code Review
di: Zeng, Zhengran, et al.
Pubblicazione: (2025)
di: Zeng, Zhengran, et al.
Pubblicazione: (2025)
A Survey of Code Review Benchmarks and Evaluation Practices in Pre-LLM and LLM Era
di: Khan, Taufiqul Islam, et al.
Pubblicazione: (2026)
di: Khan, Taufiqul Islam, et al.
Pubblicazione: (2026)
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
di: Cui, Yi
Pubblicazione: (2025)
di: Cui, Yi
Pubblicazione: (2025)
CodeArena: A Collective Evaluation Platform for LLM Code Generation
di: Du, Mingzhe, et al.
Pubblicazione: (2025)
di: Du, Mingzhe, et al.
Pubblicazione: (2025)
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
di: Farchi, Eitan, et al.
Pubblicazione: (2024)
di: Farchi, Eitan, et al.
Pubblicazione: (2024)
LLM-Based Test-Driven Interactive Code Generation: User Study and Empirical Evaluation
di: Fakhoury, Sarah, et al.
Pubblicazione: (2024)
di: Fakhoury, Sarah, et al.
Pubblicazione: (2024)
The Fault in our Stars: Quality Assessment of Code Generation Benchmarks
di: Siddiq, Mohammed Latif, et al.
Pubblicazione: (2024)
di: Siddiq, Mohammed Latif, et al.
Pubblicazione: (2024)
Comparing Developer and LLM Biases in Code Evaluation
di: Mittal, Aditya, et al.
Pubblicazione: (2026)
di: Mittal, Aditya, et al.
Pubblicazione: (2026)
DSL or Code? Evaluating the Quality of LLM-Generated Algebraic Specifications: A Case Study in Optimization at Kinaxis
di: Ayoughi, Negin, et al.
Pubblicazione: (2026)
di: Ayoughi, Negin, et al.
Pubblicazione: (2026)
COFFE: A Code Efficiency Benchmark for Code Generation
di: Peng, Yun, et al.
Pubblicazione: (2025)
di: Peng, Yun, et al.
Pubblicazione: (2025)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
di: Syromiatnikov, Mykyta, et al.
Pubblicazione: (2025)
di: Syromiatnikov, Mykyta, et al.
Pubblicazione: (2025)
Evaluating Efficiency and Novelty of LLM-Generated Code for Graph Analysis
di: Nia, Atieh Barati, et al.
Pubblicazione: (2025)
di: Nia, Atieh Barati, et al.
Pubblicazione: (2025)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
di: Pan, Zhiyuan, et al.
Pubblicazione: (2025)
di: Pan, Zhiyuan, et al.
Pubblicazione: (2025)
Continuous Benchmark Generation for Evaluating Enterprise-scale LLM Agents
di: Saxena, Divyanshu, et al.
Pubblicazione: (2025)
di: Saxena, Divyanshu, et al.
Pubblicazione: (2025)
Human or LLM? A Comparative Study on Accessible Code Generation Capability
di: Suh, Hyunjae, et al.
Pubblicazione: (2025)
di: Suh, Hyunjae, et al.
Pubblicazione: (2025)
Development and Benchmarking of Multilingual Code Clone Detector
di: Zhu, Wenqing, et al.
Pubblicazione: (2024)
di: Zhu, Wenqing, et al.
Pubblicazione: (2024)
SolContractEval: A Benchmark for Evaluating Contract-Level Solidity Code Generation
di: Ye, Zhifan, et al.
Pubblicazione: (2025)
di: Ye, Zhifan, et al.
Pubblicazione: (2025)
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations
di: Tóth, Rebeka, et al.
Pubblicazione: (2024)
di: Tóth, Rebeka, et al.
Pubblicazione: (2024)
HumanEvalComm: Benchmarking the Communication Competence of Code Generation for LLMs and LLM Agent
di: Wu, Jie JW, et al.
Pubblicazione: (2024)
di: Wu, Jie JW, et al.
Pubblicazione: (2024)
Beyond Code Similarity: Benchmarking the Plausibility, Efficiency, and Complexity of LLM-Generated Smart Contracts
di: Salzano, Francesco, et al.
Pubblicazione: (2025)
di: Salzano, Francesco, et al.
Pubblicazione: (2025)
Assessing Small Language Models for Code Generation: An Empirical Study with Benchmarks
di: Hasan, Md Mahade, et al.
Pubblicazione: (2025)
di: Hasan, Md Mahade, et al.
Pubblicazione: (2025)
ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex Code
di: Feng, Jia, et al.
Pubblicazione: (2024)
di: Feng, Jia, et al.
Pubblicazione: (2024)
Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development
di: Zeng, Zhengran, et al.
Pubblicazione: (2025)
di: Zeng, Zhengran, et al.
Pubblicazione: (2025)
Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review
di: Paul, Debalina Ghosh, et al.
Pubblicazione: (2024)
di: Paul, Debalina Ghosh, et al.
Pubblicazione: (2024)
A Differential Fuzzing-Based Evaluation of Functional Equivalence in LLM-Generated Code Refactorings
di: Dristi, Simantika Bhattacharjee, et al.
Pubblicazione: (2026)
di: Dristi, Simantika Bhattacharjee, et al.
Pubblicazione: (2026)
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation
di: Zhang, Binquan, et al.
Pubblicazione: (2025)
di: Zhang, Binquan, et al.
Pubblicazione: (2025)
SemGuard: Real-Time Semantic Evaluator for Correcting LLM-Generated Code
di: Wang, Qinglin, et al.
Pubblicazione: (2025)
di: Wang, Qinglin, et al.
Pubblicazione: (2025)
Cross-Task Benchmarking and Evaluation of General-Purpose and Code-Specific Large Language Models
di: Das, Gunjan, et al.
Pubblicazione: (2025)
di: Das, Gunjan, et al.
Pubblicazione: (2025)
A Performance Study of LLM-Generated Code on Leetcode
di: Coignion, Tristan, et al.
Pubblicazione: (2024)
di: Coignion, Tristan, et al.
Pubblicazione: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
di: Li, Jia, et al.
Pubblicazione: (2024)
di: Li, Jia, et al.
Pubblicazione: (2024)
Inducing Vulnerable Code Generation in LLM Coding Assistants
di: Zeng, Binqi, et al.
Pubblicazione: (2025)
di: Zeng, Binqi, et al.
Pubblicazione: (2025)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
di: Li, Jia, et al.
Pubblicazione: (2026)
di: Li, Jia, et al.
Pubblicazione: (2026)
ScenEval: A Benchmark for Scenario-Based Evaluation of Code Generation
di: Paul, Debalina Ghosh, et al.
Pubblicazione: (2024)
di: Paul, Debalina Ghosh, et al.
Pubblicazione: (2024)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation
di: Zheng, Dewu, et al.
Pubblicazione: (2024)
di: Zheng, Dewu, et al.
Pubblicazione: (2024)
On the Effectiveness of Training Data Optimization for LLM-based Code Generation: An Empirical Study
di: Kuang, Shiqi, et al.
Pubblicazione: (2025)
di: Kuang, Shiqi, et al.
Pubblicazione: (2025)
Copilot Arena: A Platform for Code LLM Evaluation in the Wild
di: Chi, Wayne, et al.
Pubblicazione: (2025)
di: Chi, Wayne, et al.
Pubblicazione: (2025)
CodeScore: Evaluating Code Generation by Learning Code Execution
di: Dong, Yihong, et al.
Pubblicazione: (2023)
di: Dong, Yihong, et al.
Pubblicazione: (2023)
Beyond Code Generation: Assessing Code LLM Maturity with Postconditions
di: He, Fusen, et al.
Pubblicazione: (2024)
di: He, Fusen, et al.
Pubblicazione: (2024)
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
di: Diehl, Patrick, et al.
Pubblicazione: (2025)
di: Diehl, Patrick, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Benchmarking and Studying the LLM-based Code Review
di: Zeng, Zhengran, et al.
Pubblicazione: (2025) -
A Survey of Code Review Benchmarks and Evaluation Practices in Pre-LLM and LLM Era
di: Khan, Taufiqul Islam, et al.
Pubblicazione: (2026) -
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
di: Cui, Yi
Pubblicazione: (2025) -
CodeArena: A Collective Evaluation Platform for LLM Code Generation
di: Du, Mingzhe, et al.
Pubblicazione: (2025) -
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
di: Farchi, Eitan, et al.
Pubblicazione: (2024)