Benchmarking Large Language Models for ABAP Code Generation: An Empirical Study on Iterative Improvement by Compiler Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Wallraven, Stephan, Köhne, Tim, Westenberger, Hartmut, Moser, Andreas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Empirical Study on the Performance and Energy Usage of Compiled Python Code
by: Stoico, Vincenzo, et al.
Published: (2025)
by: Stoico, Vincenzo, et al.
Published: (2025)
CodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case Generation
by: Cai, Jianfeng, et al.
Published: (2026)
by: Cai, Jianfeng, et al.
Published: (2026)
Finding Compiler Bugs through Cross-Language Code Generator and Differential Testing
by: Feng, Qiong, et al.
Published: (2025)
by: Feng, Qiong, et al.
Published: (2025)
Decompiling Rust: An Empirical Study of Compiler Optimizations and Reverse Engineering Challenges
by: Zhou, Zixu
Published: (2025)
by: Zhou, Zixu
Published: (2025)
Enhancing Translation Validation of Compiler Transformations with Large Language Models
by: Wang, Yanzhao, et al.
Published: (2024)
by: Wang, Yanzhao, et al.
Published: (2024)
Iterative Refinement of Project-Level Code Context for Precise Code Generation with Compiler Feedback
by: Bi, Zhangqian, et al.
Published: (2024)
by: Bi, Zhangqian, et al.
Published: (2024)
Strengthening Programming Comprehension in Large Language Models through Code Generation
by: Ren, Xiaoning, et al.
Published: (2025)
by: Ren, Xiaoning, et al.
Published: (2025)
COBOLAssist: Analyzing and Fixing Compilation Errors for LLM-Powered COBOL Code Generation
by: Dau, Anh T. V., et al.
Published: (2026)
by: Dau, Anh T. V., et al.
Published: (2026)
A Preliminary Study of Multilingual Code Language Models for Code Generation Task Using Translated Benchmarks
by: Dandamudi, Rohit, et al.
Published: (2024)
by: Dandamudi, Rohit, et al.
Published: (2024)
COBOL-Coder: Domain-Adapted Large Language Models for COBOL Code Generation and Translation
by: Dau, Anh T. V., et al.
Published: (2026)
by: Dau, Anh T. V., et al.
Published: (2026)
Reinforcement Learning from Compiler and Language Server Feedback
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Compilation Quotient (CQ): A Metric for the Compilation Hardness of Programming Languages
by: Szabo, Violet, et al.
Published: (2024)
by: Szabo, Violet, et al.
Published: (2024)
From Effectiveness to Efficiency: Uncovering Linguistic Bias in Large Language Model-based Code Generation
by: Jiang, Weipeng, et al.
Published: (2024)
by: Jiang, Weipeng, et al.
Published: (2024)
Compilation of Commit Changes within Java Source Code Repositories
by: Schott, Stefan, et al.
Published: (2024)
by: Schott, Stefan, et al.
Published: (2024)
ViScratch: Using Large Language Models and Gameplay Videos for Automated Feedback in Scratch
by: Si, Yuan, et al.
Published: (2025)
by: Si, Yuan, et al.
Published: (2025)
iScript: A Domain-Adapted Large Language Model and Benchmark for Physical Design Tcl Script Generation
by: Xu, Ning, et al.
Published: (2026)
by: Xu, Ning, et al.
Published: (2026)
Transforming C++11 Code to C++03 to Support Legacy Compilation Environments
by: Antal, Gábor, et al.
Published: (2024)
by: Antal, Gábor, et al.
Published: (2024)
Fine-Tuning Multilingual Language Models for Code Review: An Empirical Study on Industrial C# Projects
by: Begolli, Igli, et al.
Published: (2025)
by: Begolli, Igli, et al.
Published: (2025)
LLMigrate: Transforming "Lazy" Large Language Models into Efficient Source Code Migrators
by: Liu, Yuchen, et al.
Published: (2025)
by: Liu, Yuchen, et al.
Published: (2025)
Scalable, Validated Code Translation of Entire Projects using Large Language Models
by: Zhang, Hanliang, et al.
Published: (2024)
by: Zhang, Hanliang, et al.
Published: (2024)
HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
by: Peng, Qiwei, et al.
Published: (2024)
by: Peng, Qiwei, et al.
Published: (2024)
Accurate Coverage Metrics for Compiler-Generated Debugging Information
by: Stinnett, J. Ryan, et al.
Published: (2024)
by: Stinnett, J. Ryan, et al.
Published: (2024)
WhiteFox: White-Box Compiler Fuzzing Empowered by Large Language Models
by: Yang, Chenyuan, et al.
Published: (2023)
by: Yang, Chenyuan, et al.
Published: (2023)
A Roadmap for Tamed Interactions with Large Language Models
by: Scotti, Vincenzo, et al.
Published: (2025)
by: Scotti, Vincenzo, et al.
Published: (2025)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
by: Wang, Peiding, et al.
Published: (2025)
by: Wang, Peiding, et al.
Published: (2025)
SIMCOPILOT: Evaluating Large Language Models for Copilot-Style Code Generation
by: Jiang, Mingchao, et al.
Published: (2025)
by: Jiang, Mingchao, et al.
Published: (2025)
Semantic Source Code Segmentation using Small and Large Language Models
by: Dahou, Abdelhalim, et al.
Published: (2025)
by: Dahou, Abdelhalim, et al.
Published: (2025)
MLIR-Smith: A Novel Random Program Generator for Evaluating Compiler Pipelines
by: Ates, Berke, et al.
Published: (2026)
by: Ates, Berke, et al.
Published: (2026)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
by: Cao, Jialun, et al.
Published: (2024)
by: Cao, Jialun, et al.
Published: (2024)
Fully Automated Generation of Combinatorial Optimisation Systems Using Large Language Models
by: Karapetyan, Daniel
Published: (2025)
by: Karapetyan, Daniel
Published: (2025)
AI-Mediated Code Comment Improvement
by: Dhakal, Maria, et al.
Published: (2025)
by: Dhakal, Maria, et al.
Published: (2025)
CodeMind: Evaluating Large Language Models for Code Reasoning
by: Liu, Changshu, et al.
Published: (2024)
by: Liu, Changshu, et al.
Published: (2024)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
by: Zhang, William, et al.
Published: (2024)
by: Zhang, William, et al.
Published: (2024)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
by: Peng, Yun, et al.
Published: (2024)
by: Peng, Yun, et al.
Published: (2024)
Perish or Flourish? A Holistic Evaluation of Large Language Models for Code Generation in Functional Programming
by: Lang, Nguyet-Anh H., et al.
Published: (2026)
by: Lang, Nguyet-Anh H., et al.
Published: (2026)
Language Models for Code Completion: A Practical Evaluation
by: Izadi, Maliheh, et al.
Published: (2024)
by: Izadi, Maliheh, et al.
Published: (2024)
Finding Missed Code Size Optimizations in Compilers using LLMs
by: Italiano, Davide, et al.
Published: (2024)
by: Italiano, Davide, et al.
Published: (2024)
Bootstrapping Fuzzers for Compilers of Low-Resource Language Dialects Using Language Models
by: Vaidya, Sairam, et al.
Published: (2025)
by: Vaidya, Sairam, et al.
Published: (2025)
Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny
by: Carreira, Carolina, et al.
Published: (2025)
by: Carreira, Carolina, et al.
Published: (2025)
DeCon: Detecting Incorrect Assertions via Postconditions Generated by a Large Language Model
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
Similar Items
-
An Empirical Study on the Performance and Energy Usage of Compiled Python Code
by: Stoico, Vincenzo, et al.
Published: (2025) -
CodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case Generation
by: Cai, Jianfeng, et al.
Published: (2026) -
Finding Compiler Bugs through Cross-Language Code Generator and Differential Testing
by: Feng, Qiong, et al.
Published: (2025) -
Decompiling Rust: An Empirical Study of Compiler Optimizations and Reverse Engineering Challenges
by: Zhou, Zixu
Published: (2025) -
Enhancing Translation Validation of Compiler Transformations with Large Language Models
by: Wang, Yanzhao, et al.
Published: (2024)