LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Diehl, Patrick, Nader, Nojoud, Moraru, Maxim, Brandt, Steven R.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912470751248384
author Diehl, Patrick
Nader, Nojoud
Moraru, Maxim
Brandt, Steven R.
author_facet Diehl, Patrick
Nader, Nojoud
Moraru, Maxim
Brandt, Steven R.
contents The rapid evolution of large language models (LLMs) has opened new possibilities for automating various tasks in software development. This paper evaluates the capabilities of the Llama 2-70B model in automating these tasks for scientific applications written in commonly used programming languages. Using representative test problems, we assess the model's capacity to generate code, documentation, and unit tests, as well as its ability to translate existing code between commonly used programming languages. Our comprehensive analysis evaluates the compilation, runtime behavior, and correctness of the generated and translated code. Additionally, we assess the quality of automatically generated code, documentation and unit tests. Our results indicate that while Llama 2-70B frequently generates syntactically correct and functional code for simpler numerical tasks, it encounters substantial difficulties with more complex, parallelized, or distributed computations, requiring considerable manual corrections. We identify key limitations and suggest areas for future improvements to better leverage AI-driven automation in scientific computing workflows.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19217
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
Diehl, Patrick
Nader, Nojoud
Moraru, Maxim
Brandt, Steven R.
Software Engineering
Artificial Intelligence
Machine Learning
The rapid evolution of large language models (LLMs) has opened new possibilities for automating various tasks in software development. This paper evaluates the capabilities of the Llama 2-70B model in automating these tasks for scientific applications written in commonly used programming languages. Using representative test problems, we assess the model's capacity to generate code, documentation, and unit tests, as well as its ability to translate existing code between commonly used programming languages. Our comprehensive analysis evaluates the compilation, runtime behavior, and correctness of the generated and translated code. Additionally, we assess the quality of automatically generated code, documentation and unit tests. Our results indicate that while Llama 2-70B frequently generates syntactically correct and functional code for simpler numerical tasks, it encounters substantial difficulties with more complex, parallelized, or distributed computations, requiring considerable manual corrections. We identify key limitations and suggest areas for future improvements to better leverage AI-driven automation in scientific computing workflows.
title LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
topic Software Engineering
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.19217