Evaluating AI-generated code for C++, Fortran, Go, Java, Julia, Matlab, Python, R, and Rust
Fuente:
arXiv
Saved in:
| Main Authors: | Diehl, Patrick, Nader, Noujoud, Brandt, Steve, Kaiser, Hartmut |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can LLMs Find Bugs in Code? An Evaluation from Beginner Errors to Security Vulnerabilities in Python and C++
by: Mhatre, Akshay, et al.
Published: (2025)
by: Mhatre, Akshay, et al.
Published: (2025)
LLM & HPC:Benchmarking DeepSeek's Performance in High-Performance Computing Tasks
by: Nader, Noujoud, et al.
Published: (2025)
by: Nader, Noujoud, et al.
Published: (2025)
LLM-HPC++: Evaluating LLM-Generated Modern C++ and MPI+OpenMP Codes for Scalable Mandelbrot Set Computation
by: Diehl, Patrick, et al.
Published: (2025)
by: Diehl, Patrick, et al.
Published: (2025)
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
by: Diehl, Patrick, et al.
Published: (2025)
by: Diehl, Patrick, et al.
Published: (2025)
From Legacy Fortran to Portable Kokkos: An Autonomous Agentic AI Workflow
by: Gupta, Sparsh, et al.
Published: (2025)
by: Gupta, Sparsh, et al.
Published: (2025)
From Translation to Superset: Benchmark-Driven Evolution of a Production AI Agent from Rust to Python
by: Wang, Jinhua, et al.
Published: (2026)
by: Wang, Jinhua, et al.
Published: (2026)
REMODEL-LLM: Transforming C code to Java using LLMs
by: Gupta, Aryan, et al.
Published: (2025)
by: Gupta, Aryan, et al.
Published: (2025)
EvoC2Rust: A Skeleton-guided Framework for Project-Level C-to-Rust Translation
by: Wang, Chaofan, et al.
Published: (2025)
by: Wang, Chaofan, et al.
Published: (2025)
FreshBrew: A Benchmark for Evaluating AI Agents on Java Code Migration
by: May, Victor, et al.
Published: (2025)
by: May, Victor, et al.
Published: (2025)
Project-Level C-to-Rust Translation via Pointer Knowledge Graphs
by: Yuan, Zhiqiang, et al.
Published: (2025)
by: Yuan, Zhiqiang, et al.
Published: (2025)
Quality Evaluation of COBOL to Java Code Transformation
by: Froimovich, Shmulik, et al.
Published: (2025)
by: Froimovich, Shmulik, et al.
Published: (2025)
RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation
by: Liang, Linxi, et al.
Published: (2025)
by: Liang, Linxi, et al.
Published: (2025)
Evaluating perturbation robustness of generative systems that use COBOL code inputs
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
Benchmarking the Parallel 1D Heat Equation Solver in Chapel, Charm++, C++, HPX, Go, Julia, Python, Rust, Swift, and Java
by: Diehl, Patrick, et al.
Published: (2023)
by: Diehl, Patrick, et al.
Published: (2023)
Quality and Security Signals in AI-Generated Python Refactoring Pull Requests
by: Almukhtar, Mohamed, et al.
Published: (2026)
by: Almukhtar, Mohamed, et al.
Published: (2026)
Automated Proof Generation for Rust Code via Self-Evolution
by: Chen, Tianyu, et al.
Published: (2024)
by: Chen, Tianyu, et al.
Published: (2024)
Adaptive Hierarchical Evaluation of LLMs and SAST tools for CWE Prediction in Python
by: Adnan, Muntasir, et al.
Published: (2026)
by: Adnan, Muntasir, et al.
Published: (2026)
PyGen: A Collaborative Human-AI Approach to Python Package Creation
by: Barua, Saikat, et al.
Published: (2024)
by: Barua, Saikat, et al.
Published: (2024)
Developing a High-Performance Process Mining Library with Java and Python Bindings in Rust
by: Küsters, Aaron, et al.
Published: (2024)
by: Küsters, Aaron, et al.
Published: (2024)
Raw Pointer Rewriting with LLMs for Translating C to Safer Rust
by: Gao, Yifei, et al.
Published: (2025)
by: Gao, Yifei, et al.
Published: (2025)
Feedback Loops and Code Perturbations in LLM-based Software Engineering: A Case Study on a C-to-Rust Translation System
by: Weiss, Martin, et al.
Published: (2025)
by: Weiss, Martin, et al.
Published: (2025)
Go-Oracle: Automated Test Oracle for Go Concurrency Bugs
by: Tsimpourlas, Foivos, et al.
Published: (2024)
by: Tsimpourlas, Foivos, et al.
Published: (2024)
Automated Testing of COBOL to Java Transformation
by: Hans, Sandeep, et al.
Published: (2025)
by: Hans, Sandeep, et al.
Published: (2025)
Automated Validation of COBOL to Java Transformation
by: Kumar, Atul, et al.
Published: (2025)
by: Kumar, Atul, et al.
Published: (2025)
Assessing LLM code generation quality through path planning tasks
by: Chen, Wanyi, et al.
Published: (2025)
by: Chen, Wanyi, et al.
Published: (2025)
Lifecycle-Aware code generation: Leveraging Software Engineering Phases in LLMs
by: Xing, Xing, et al.
Published: (2025)
by: Xing, Xing, et al.
Published: (2025)
Canonical Intermediate Representation for LLM-based optimization problem formulation and code generation
by: Lyu, Zhongyuan, et al.
Published: (2026)
by: Lyu, Zhongyuan, et al.
Published: (2026)
Embedding Software Intent: Lightweight Java Module Recovery
by: He, Yirui, et al.
Published: (2025)
by: He, Yirui, et al.
Published: (2025)
PyVeritas: On Verifying Python via LLM-Based Transpilation and Bounded Model Checking for C
by: Orvalho, Pedro, et al.
Published: (2025)
by: Orvalho, Pedro, et al.
Published: (2025)
Automated QoR improvement in OpenROAD with coding agents
by: Ghose, Amur, et al.
Published: (2026)
by: Ghose, Amur, et al.
Published: (2026)
When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions
by: Larbi, Maya, et al.
Published: (2025)
by: Larbi, Maya, et al.
Published: (2025)
GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities
by: Misra, Diganta, et al.
Published: (2025)
by: Misra, Diganta, et al.
Published: (2025)
Better Python Programming for all: With the focus on Maintainability
by: Shivashankar, Karthik, et al.
Published: (2024)
by: Shivashankar, Karthik, et al.
Published: (2024)
Call-Chain-Aware LLM-Based Test Generation for Java Projects
by: Wang, Guancheng, et al.
Published: (2026)
by: Wang, Guancheng, et al.
Published: (2026)
A systematic review of generative AI usage for IT project management
by: Anghel, Ionut, et al.
Published: (2026)
by: Anghel, Ionut, et al.
Published: (2026)
Automating the Correctness Assessment of AI-generated Code for Security Contexts
by: Cotroneo, Domenico, et al.
Published: (2023)
by: Cotroneo, Domenico, et al.
Published: (2023)
ENCRUST: Encapsulated Substitution and Agentic Refinement on a Live Scaffold for Safe C-to-Rust Translation
by: Sim, Hohyun, et al.
Published: (2026)
by: Sim, Hohyun, et al.
Published: (2026)
GoNoGo: An Efficient LLM-based Multi-Agent System for Streamlining Automotive Software Release Decision-Making
by: Khoee, Arsham Gholamzadeh, et al.
Published: (2024)
by: Khoee, Arsham Gholamzadeh, et al.
Published: (2024)
LLMs for Automated Unit Test Generation and Assessment in Java: The AgoneTest Framework
by: Lops, Andrea, et al.
Published: (2025)
by: Lops, Andrea, et al.
Published: (2025)
GenAI-based test case generation and execution in SDV platform
by: Zyberaj, Denesa, et al.
Published: (2025)
by: Zyberaj, Denesa, et al.
Published: (2025)
Similar Items
-
Can LLMs Find Bugs in Code? An Evaluation from Beginner Errors to Security Vulnerabilities in Python and C++
by: Mhatre, Akshay, et al.
Published: (2025) -
LLM & HPC:Benchmarking DeepSeek's Performance in High-Performance Computing Tasks
by: Nader, Noujoud, et al.
Published: (2025) -
LLM-HPC++: Evaluating LLM-Generated Modern C++ and MPI+OpenMP Codes for Scalable Mandelbrot Set Computation
by: Diehl, Patrick, et al.
Published: (2025) -
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
by: Diehl, Patrick, et al.
Published: (2025) -
From Legacy Fortran to Portable Kokkos: An Autonomous Agentic AI Workflow
by: Gupta, Sparsh, et al.
Published: (2025)