PPM: Automated Generation of Diverse Programming Problems for Benchmarking Code Generation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Simin, Feng, Xiaoning, Han, Xiaohong, Liu, Cong, Yang, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLMEffiChecker: Understanding and Testing Efficiency Degradation of Large Language Models
von: Feng, Xiaoning, et al.
Veröffentlicht: (2022)
von: Feng, Xiaoning, et al.
Veröffentlicht: (2022)
AI Coders Are Among Us: Rethinking Programming Language Grammar Towards Efficient Code Generation
von: Sun, Zhensu, et al.
Veröffentlicht: (2024)
von: Sun, Zhensu, et al.
Veröffentlicht: (2024)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
von: Zhang, William, et al.
Veröffentlicht: (2024)
von: Zhang, William, et al.
Veröffentlicht: (2024)
Enhancing Automated Loop Invariant Generation for Complex Programs with Large Language Models
von: Liu, Ruibang, et al.
Veröffentlicht: (2024)
von: Liu, Ruibang, et al.
Veröffentlicht: (2024)
Perish or Flourish? A Holistic Evaluation of Large Language Models for Code Generation in Functional Programming
von: Lang, Nguyet-Anh H., et al.
Veröffentlicht: (2026)
von: Lang, Nguyet-Anh H., et al.
Veröffentlicht: (2026)
Is Functional Correctness Enough to Evaluate Code Language Models? Exploring Diversity of Generated Codes
von: Chon, Heejae, et al.
Veröffentlicht: (2024)
von: Chon, Heejae, et al.
Veröffentlicht: (2024)
AutoCode: LLMs as Problem Setters for Competitive Programming
von: Zhou, Shang, et al.
Veröffentlicht: (2025)
von: Zhou, Shang, et al.
Veröffentlicht: (2025)
A Preliminary Study of Multilingual Code Language Models for Code Generation Task Using Translated Benchmarks
von: Dandamudi, Rohit, et al.
Veröffentlicht: (2024)
von: Dandamudi, Rohit, et al.
Veröffentlicht: (2024)
Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation
von: Fang, Sen, et al.
Veröffentlicht: (2025)
von: Fang, Sen, et al.
Veröffentlicht: (2025)
Executing as You Generate: Hiding Execution Latency in LLM Code Generation
von: Sun, Zhensu, et al.
Veröffentlicht: (2026)
von: Sun, Zhensu, et al.
Veröffentlicht: (2026)
Benchmarking Large Language Models for ABAP Code Generation: An Empirical Study on Iterative Improvement by Compiler Feedback
von: Wallraven, Stephan, et al.
Veröffentlicht: (2026)
von: Wallraven, Stephan, et al.
Veröffentlicht: (2026)
CodeMind: Evaluating Large Language Models for Code Reasoning
von: Liu, Changshu, et al.
Veröffentlicht: (2024)
von: Liu, Changshu, et al.
Veröffentlicht: (2024)
Learning to Guarantee Type Correctness in Code Generation through Type-Guided Program Synthesis
von: Huang, Zhechong, et al.
Veröffentlicht: (2025)
von: Huang, Zhechong, et al.
Veröffentlicht: (2025)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
von: Wang, Peiding, et al.
Veröffentlicht: (2025)
von: Wang, Peiding, et al.
Veröffentlicht: (2025)
Self-Improving Code Generation via Semantic Entropy and Behavioral Consensus
von: Zhang, Huan, et al.
Veröffentlicht: (2026)
von: Zhang, Huan, et al.
Veröffentlicht: (2026)
Strengthening Programming Comprehension in Large Language Models through Code Generation
von: Ren, Xiaoning, et al.
Veröffentlicht: (2025)
von: Ren, Xiaoning, et al.
Veröffentlicht: (2025)
Program Skeletons for Automated Program Translation
von: Wang, Bo, et al.
Veröffentlicht: (2025)
von: Wang, Bo, et al.
Veröffentlicht: (2025)
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation
von: Chen, Le, et al.
Veröffentlicht: (2025)
von: Chen, Le, et al.
Veröffentlicht: (2025)
From Code Generation to Software Testing: AI Copilot with Context-Based RAG
von: Wang, Yuchen, et al.
Veröffentlicht: (2025)
von: Wang, Yuchen, et al.
Veröffentlicht: (2025)
Dynamic Stability of LLM-Generated Code
von: Rajput, Prateek, et al.
Veröffentlicht: (2025)
von: Rajput, Prateek, et al.
Veröffentlicht: (2025)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
von: Peng, Yun, et al.
Veröffentlicht: (2024)
von: Peng, Yun, et al.
Veröffentlicht: (2024)
From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging
von: Shi, Yuling, et al.
Veröffentlicht: (2024)
von: Shi, Yuling, et al.
Veröffentlicht: (2024)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
von: Chen, Simin, et al.
Veröffentlicht: (2025)
von: Chen, Simin, et al.
Veröffentlicht: (2025)
Effective LLM-Driven Code Generation with Pythoness
von: Levin, Kyla H., et al.
Veröffentlicht: (2025)
von: Levin, Kyla H., et al.
Veröffentlicht: (2025)
Code Broker: A Multi-Agent System for Automated Code Quality Assessment
von: Attrah, Samer
Veröffentlicht: (2026)
von: Attrah, Samer
Veröffentlicht: (2026)
Is Self-Repair a Silver Bullet for Code Generation?
von: Olausson, Theo X., et al.
Veröffentlicht: (2023)
von: Olausson, Theo X., et al.
Veröffentlicht: (2023)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
von: Cao, Jialun, et al.
Veröffentlicht: (2024)
von: Cao, Jialun, et al.
Veröffentlicht: (2024)
Ranking LLM-Generated Loop Invariants for Program Verification
von: Chakraborty, Saikat, et al.
Veröffentlicht: (2023)
von: Chakraborty, Saikat, et al.
Veröffentlicht: (2023)
How Diversely Can Language Models Solve Problems? Exploring the Algorithmic Diversity of Model-Generated Code
von: Lee, Seonghyeon, et al.
Veröffentlicht: (2025)
von: Lee, Seonghyeon, et al.
Veröffentlicht: (2025)
Hydra: Efficient, Correct Code Generation via Checkpoint-and-Rollback Support
von: Du, Alexander, et al.
Veröffentlicht: (2026)
von: Du, Alexander, et al.
Veröffentlicht: (2026)
Assessing GPT-4-Vision's Capabilities in UML-Based Code Generation
von: Antal, Gábor, et al.
Veröffentlicht: (2024)
von: Antal, Gábor, et al.
Veröffentlicht: (2024)
A Problem-Oriented Perspective and Anchor Verification for Code Optimization
von: Ye, Tong, et al.
Veröffentlicht: (2024)
von: Ye, Tong, et al.
Veröffentlicht: (2024)
ACCeLLiuM: Supervised Fine-Tuning for Automated OpenACC Pragma Generation
von: Jhaveri, Samyak, et al.
Veröffentlicht: (2025)
von: Jhaveri, Samyak, et al.
Veröffentlicht: (2025)
AutoMCQ -- Automatically Generate Code Comprehension Questions using GenAI
von: Goodfellow, Martin, et al.
Veröffentlicht: (2025)
von: Goodfellow, Martin, et al.
Veröffentlicht: (2025)
IndustryCode: A Benchmark for Industry Code Generation
von: Zeng, Puyu, et al.
Veröffentlicht: (2026)
von: Zeng, Puyu, et al.
Veröffentlicht: (2026)
Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2025)
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2025)
EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
ECO: Enhanced Code Optimization via Performance-Aware Prompting for Code-LLMs
von: Kim, Su-Hyeon, et al.
Veröffentlicht: (2025)
von: Kim, Su-Hyeon, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLMEffiChecker: Understanding and Testing Efficiency Degradation of Large Language Models
von: Feng, Xiaoning, et al.
Veröffentlicht: (2022) -
AI Coders Are Among Us: Rethinking Programming Language Grammar Towards Efficient Code Generation
von: Sun, Zhensu, et al.
Veröffentlicht: (2024) -
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
von: Zhang, William, et al.
Veröffentlicht: (2024) -
Enhancing Automated Loop Invariant Generation for Complex Programs with Large Language Models
von: Liu, Ruibang, et al.
Veröffentlicht: (2024) -
Perish or Flourish? A Holistic Evaluation of Large Language Models for Code Generation in Functional Programming
von: Lang, Nguyet-Anh H., et al.
Veröffentlicht: (2026)