LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Diehl, Patrick, Nader, Nojoud, Moraru, Maxim, Brandt, Steven R. |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
CUBETESTERAI: Automated JUnit Test Generation using the LLaMA Model
par: Gorla, Daniele, et autres
Publié: (2025)
par: Gorla, Daniele, et autres
Publié: (2025)
Evaluating LLaMA 3.2 for Software Vulnerability Detection
par: Gonçalves, José, et autres
Publié: (2025)
par: Gonçalves, José, et autres
Publié: (2025)
Evaluating AI-generated code for C++, Fortran, Go, Java, Julia, Matlab, Python, R, and Rust
par: Diehl, Patrick, et autres
Publié: (2024)
par: Diehl, Patrick, et autres
Publié: (2024)
Smart-LLaMA-DPO: Reinforced Large Language Model for Explainable Smart Contract Vulnerability Detection
par: Yu, Lei, et autres
Publié: (2025)
par: Yu, Lei, et autres
Publié: (2025)
Smart-LLaMA: Two-Stage Post-Training of Large Language Models for Smart Contract Vulnerability Detection and Explanation
par: Yu, Lei, et autres
Publié: (2024)
par: Yu, Lei, et autres
Publié: (2024)
LLM-HPC++: Evaluating LLM-Generated Modern C++ and MPI+OpenMP Codes for Scalable Mandelbrot Set Computation
par: Diehl, Patrick, et autres
Publié: (2025)
par: Diehl, Patrick, et autres
Publié: (2025)
From Legacy Fortran to Portable Kokkos: An Autonomous Agentic AI Workflow
par: Gupta, Sparsh, et autres
Publié: (2025)
par: Gupta, Sparsh, et autres
Publié: (2025)
Can LLMs Find Bugs in Code? An Evaluation from Beginner Errors to Security Vulnerabilities in Python and C++
par: Mhatre, Akshay, et autres
Publié: (2025)
par: Mhatre, Akshay, et autres
Publié: (2025)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
par: Pan, Zhiyuan, et autres
Publié: (2025)
par: Pan, Zhiyuan, et autres
Publié: (2025)
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
par: Cui, Yi
Publié: (2025)
par: Cui, Yi
Publié: (2025)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
par: Zhang, Wentao, et autres
Publié: (2026)
par: Zhang, Wentao, et autres
Publié: (2026)
How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks
par: Arimbur, Johin Johny
Publié: (2026)
par: Arimbur, Johin Johny
Publié: (2026)
CrossPL: Evaluating Large Language Models on Cross Programming Language Code Generation
par: Xiong, Zhanhang, et autres
Publié: (2025)
par: Xiong, Zhanhang, et autres
Publié: (2025)
MLDebugging: Towards Benchmarking Code Debugging Across Multi-Library Scenarios
par: Huang, Jinyang, et autres
Publié: (2025)
par: Huang, Jinyang, et autres
Publié: (2025)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
par: Rahman, Musfiqur, et autres
Publié: (2025)
par: Rahman, Musfiqur, et autres
Publié: (2025)
CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models
par: Padwal, Vedant
Publié: (2026)
par: Padwal, Vedant
Publié: (2026)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
par: Zhang, William, et autres
Publié: (2024)
par: Zhang, William, et autres
Publié: (2024)
Evaluating LLM-Generated Code: A Benchmark and Developer Study
par: Szych, Joanna, et autres
Publié: (2026)
par: Szych, Joanna, et autres
Publié: (2026)
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations
par: Tóth, Rebeka, et autres
Publié: (2024)
par: Tóth, Rebeka, et autres
Publié: (2024)
COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models
par: Meaden, James, et autres
Publié: (2025)
par: Meaden, James, et autres
Publié: (2025)
AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
par: Zhang, Tanghaoran, et autres
Publié: (2026)
par: Zhang, Tanghaoran, et autres
Publié: (2026)
CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning
par: Lam, Man Ho, et autres
Publié: (2025)
par: Lam, Man Ho, et autres
Publié: (2025)
Energy-Aware Code Generation with LLMs: Benchmarking Small vs. Large Language Models for Sustainable AI Programming
par: Ashraf, Humza, et autres
Publié: (2025)
par: Ashraf, Humza, et autres
Publié: (2025)
RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair
par: Silva, André, et autres
Publié: (2023)
par: Silva, André, et autres
Publié: (2023)
Prompt Driven Development with Claude Code: Building a Complete TUI Framework for the Ring Programming Language
par: Fayed, Mahmoud Samir, et autres
Publié: (2026)
par: Fayed, Mahmoud Samir, et autres
Publié: (2026)
Software Development Life Cycle Perspective: A Survey of Benchmarks for Code Large Language Models and Agents
par: Wang, Kaixin, et autres
Publié: (2025)
par: Wang, Kaixin, et autres
Publié: (2025)
Cracking CodeWhisperer: Analyzing Developers' Interactions and Patterns During Programming Tasks
par: Javahar, Jeena, et autres
Publié: (2025)
par: Javahar, Jeena, et autres
Publié: (2025)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
par: Zhou, Qixing, et autres
Publié: (2026)
par: Zhou, Qixing, et autres
Publié: (2026)
A Performance Study of LLM-Generated Code on Leetcode
par: Coignion, Tristan, et autres
Publié: (2024)
par: Coignion, Tristan, et autres
Publié: (2024)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
par: Agarwal, Anisha, et autres
Publié: (2024)
par: Agarwal, Anisha, et autres
Publié: (2024)
Evaluation of the Code Generation Capabilities of ChatGPT 4: A Comparative Analysis in 19 Programming Languages
par: Gilbert, L. C.
Publié: (2025)
par: Gilbert, L. C.
Publié: (2025)
CoCo-Bench: A Comprehensive Code Benchmark For Multi-task Large Language Model Evaluation
par: Yin, Wenjing, et autres
Publié: (2025)
par: Yin, Wenjing, et autres
Publié: (2025)
LLM Code Customization with Visual Results: A Benchmark on TikZ
par: Reux, Charly, et autres
Publié: (2025)
par: Reux, Charly, et autres
Publié: (2025)
A New Benchmark for the Appropriate Evaluation of RTL Code Optimization
par: Lu, Yao, et autres
Publié: (2026)
par: Lu, Yao, et autres
Publié: (2026)
Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review
par: Paul, Debalina Ghosh, et autres
Publié: (2024)
par: Paul, Debalina Ghosh, et autres
Publié: (2024)
Automating Structural Analysis Across Multiple Software Platforms Using Large Language Models
par: Geng, Ziheng, et autres
Publié: (2026)
par: Geng, Ziheng, et autres
Publié: (2026)
CodeFuse-CommitEval: Towards Benchmarking LLM's Power on Commit Message and Code Change Inconsistency Detection
par: Zhang, Qingyu, et autres
Publié: (2025)
par: Zhang, Qingyu, et autres
Publié: (2025)
Do Current Language Models Support Code Intelligence for R Programming Language?
par: Zhao, ZiXiao, et autres
Publié: (2024)
par: Zhao, ZiXiao, et autres
Publié: (2024)
An Execution-Verified Multi-Language Benchmark for Code Semantic Reasoning
par: Li, Yikun, et autres
Publié: (2026)
par: Li, Yikun, et autres
Publié: (2026)
Documents similaires
-
CUBETESTERAI: Automated JUnit Test Generation using the LLaMA Model
par: Gorla, Daniele, et autres
Publié: (2025) -
Evaluating LLaMA 3.2 for Software Vulnerability Detection
par: Gonçalves, José, et autres
Publié: (2025) -
Evaluating AI-generated code for C++, Fortran, Go, Java, Julia, Matlab, Python, R, and Rust
par: Diehl, Patrick, et autres
Publié: (2024) -
Smart-LLaMA-DPO: Reinforced Large Language Model for Explainable Smart Contract Vulnerability Detection
par: Yu, Lei, et autres
Publié: (2025) -
Smart-LLaMA: Two-Stage Post-Training of Large Language Models for Smart Contract Vulnerability Detection and Explanation
par: Yu, Lei, et autres
Publié: (2024)