Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Farchi, Eitan, Froimovich, Shmulik, Katan, Rami, Raz, Orna |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
PACIFIC: a framework for generating benchmarks to check Precise Automatically Checked Instruction Following In Code
by: Dreyfuss, Itay, et al.
Published: (2025)
by: Dreyfuss, Itay, et al.
Published: (2025)
Quality Evaluation of COBOL to Java Code Transformation
by: Froimovich, Shmulik, et al.
Published: (2025)
by: Froimovich, Shmulik, et al.
Published: (2025)
Using Combinatorial Optimization to Design a High quality LLM Solution
by: Ackerman, Samuel, et al.
Published: (2024)
by: Ackerman, Samuel, et al.
Published: (2024)
An Agent-Based Framework for the Automatic Validation of Mathematical Optimization Models
by: Zadorojniy, Alexander, et al.
Published: (2025)
by: Zadorojniy, Alexander, et al.
Published: (2025)
Effective Technical Reviews
by: Ballentine, Scott, et al.
Published: (2024)
by: Ballentine, Scott, et al.
Published: (2024)
A Practical Approach to Combinatorial Test Design
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
Quality Engineering for Agile and DevOps on the Cloud and Edge
by: Farchi, Eitan, et al.
Published: (2023)
by: Farchi, Eitan, et al.
Published: (2023)
Uncovering Code Insights: Leveraging GitHub Artifacts for Deeper Code Understanding
by: Nevo, Ziv, et al.
Published: (2025)
by: Nevo, Ziv, et al.
Published: (2025)
Enhancing Formal Software Specification with Artificial Intelligence
by: Nassar, Antonio Abu, et al.
Published: (2026)
by: Nassar, Antonio Abu, et al.
Published: (2026)
Generalized Coverage Criteria for Combinatorial Sequence Testing
by: Elyasaf, Achiya, et al.
Published: (2022)
by: Elyasaf, Achiya, et al.
Published: (2022)
Technique to Baseline QE Artefact Generation Aligned to Quality Metrics
by: Farchi, Eitan, et al.
Published: (2025)
by: Farchi, Eitan, et al.
Published: (2025)
Evaluating perturbation robustness of generative systems that use COBOL code inputs
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
Black-Box Bug-Amplification for Multithreaded Software
by: Weiss, Yeshayahu, et al.
Published: (2025)
by: Weiss, Yeshayahu, et al.
Published: (2025)
Is LLM-Generated Code More Maintainable \& Reliable than Human-Written Code?
by: Molison, Alfred Santa, et al.
Published: (2025)
by: Molison, Alfred Santa, et al.
Published: (2025)
Generating Unseen Code Tests In Infinitum
by: Zalmanovici, Marcel, et al.
Published: (2024)
by: Zalmanovici, Marcel, et al.
Published: (2024)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
by: Huang, Dong, et al.
Published: (2024)
by: Huang, Dong, et al.
Published: (2024)
COMCAT: Leveraging Human Judgment to Improve Automatic Documentation and Summarization
by: Grandel, Skyler, et al.
Published: (2024)
by: Grandel, Skyler, et al.
Published: (2024)
Evaluating LLM-Generated Code: A Benchmark and Developer Study
by: Szych, Joanna, et al.
Published: (2026)
by: Szych, Joanna, et al.
Published: (2026)
RepoMod-Bench: A Benchmark for Code Repository Modernization via Implementation-Agnostic Testing
by: Li, Xuefeng, et al.
Published: (2026)
by: Li, Xuefeng, et al.
Published: (2026)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
by: Chou, Jason, et al.
Published: (2025)
by: Chou, Jason, et al.
Published: (2025)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
Multi-Agent Code-Orchestrated Generation for Reliable Infrastructure-as-Code
by: Khan, Rana Nameer Hussain, et al.
Published: (2025)
by: Khan, Rana Nameer Hussain, et al.
Published: (2025)
LLM Performance for Code Generation on Noisy Tasks
by: Sendyka, Radzim, et al.
Published: (2025)
by: Sendyka, Radzim, et al.
Published: (2025)
Cross-Task Benchmarking and Evaluation of General-Purpose and Code-Specific Large Language Models
by: Das, Gunjan, et al.
Published: (2025)
by: Das, Gunjan, et al.
Published: (2025)
Automated Benchmark Generation for Repository-Level Coding Tasks
by: Vergopoulos, Konstantinos, et al.
Published: (2025)
by: Vergopoulos, Konstantinos, et al.
Published: (2025)
Benchmarking and Studying the LLM-based Code Review
by: Zeng, Zhengran, et al.
Published: (2025)
by: Zeng, Zhengran, et al.
Published: (2025)
De-Hallucinator: Mitigating LLM Hallucinations in Code Generation Tasks via Iterative Grounding
by: Eghbali, Aryaz, et al.
Published: (2024)
by: Eghbali, Aryaz, et al.
Published: (2024)
CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs
by: Manh, Dung Nguyen, et al.
Published: (2024)
by: Manh, Dung Nguyen, et al.
Published: (2024)
HumanEvalComm: Benchmarking the Communication Competence of Code Generation for LLMs and LLM Agent
by: Wu, Jie JW, et al.
Published: (2024)
by: Wu, Jie JW, et al.
Published: (2024)
Beyond Code Similarity: Benchmarking the Plausibility, Efficiency, and Complexity of LLM-Generated Smart Contracts
by: Salzano, Francesco, et al.
Published: (2025)
by: Salzano, Francesco, et al.
Published: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
by: Jiang, Hongchao, et al.
Published: (2025)
by: Jiang, Hongchao, et al.
Published: (2025)
COFFE: A Code Efficiency Benchmark for Code Generation
by: Peng, Yun, et al.
Published: (2025)
by: Peng, Yun, et al.
Published: (2025)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
by: Zhao, Songwen, et al.
Published: (2025)
by: Zhao, Songwen, et al.
Published: (2025)
Prompt Alchemy: Automatic Prompt Refinement for Enhancing Code Generation
by: Ye, Sixiang, et al.
Published: (2025)
by: Ye, Sixiang, et al.
Published: (2025)
Inducing Vulnerable Code Generation in LLM Coding Assistants
by: Zeng, Binqi, et al.
Published: (2025)
by: Zeng, Binqi, et al.
Published: (2025)
No Man is an Island: Towards Fully Automatic Programming by Code Search, Code Generation and Program Repair
by: Zhang, Quanjun, et al.
Published: (2024)
by: Zhang, Quanjun, et al.
Published: (2024)
Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency
by: Ashrafi, Nazmus, et al.
Published: (2025)
by: Ashrafi, Nazmus, et al.
Published: (2025)
Similar Items
-
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025) -
Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
by: Fandina, Ora Nova, et al.
Published: (2025) -
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
by: Fandina, Ora Nova, et al.
Published: (2025) -
PACIFIC: a framework for generating benchmarks to check Precise Automatically Checked Instruction Following In Code
by: Dreyfuss, Itay, et al.
Published: (2025) -
Quality Evaluation of COBOL to Java Code Transformation
by: Froimovich, Shmulik, et al.
Published: (2025)