Large Language Models in Code Co-generation for Safe Autonomous Vehicles

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nouri, Ali, Cabrero-Daniel, Beatriz, Fei, Zhennan, Ronanki, Krishna, Sivencrona, Håkan, Berger, Christian
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908380062285824
author Nouri, Ali
Cabrero-Daniel, Beatriz
Fei, Zhennan
Ronanki, Krishna
Sivencrona, Håkan
Berger, Christian
author_facet Nouri, Ali
Cabrero-Daniel, Beatriz
Fei, Zhennan
Ronanki, Krishna
Sivencrona, Håkan
Berger, Christian
contents Software engineers in various industrial domains are already using Large Language Models (LLMs) to accelerate the process of implementing parts of software systems. When considering its potential use for ADAS or AD systems in the automotive context, there is a need to systematically assess this new setup: LLMs entail a well-documented set of risks for safety-related systems' development due to their stochastic nature. To reduce the effort for code reviewers to evaluate LLM-generated code, we propose an evaluation pipeline to conduct sanity-checks on the generated code. We compare the performance of six state-of-the-art LLMs (CodeLlama, CodeGemma, DeepSeek-r1, DeepSeek-Coders, Mistral, and GPT-4) on four safety-related programming tasks. Additionally, we qualitatively analyse the most frequent faults generated by these LLMs, creating a failure-mode catalogue to support human reviewers. Finally, the limitations and capabilities of LLMs in code generation, and the use of the proposed pipeline in the existing process, are discussed.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19658
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Models in Code Co-generation for Safe Autonomous Vehicles
Nouri, Ali
Cabrero-Daniel, Beatriz
Fei, Zhennan
Ronanki, Krishna
Sivencrona, Håkan
Berger, Christian
Software Engineering
Artificial Intelligence
Software engineers in various industrial domains are already using Large Language Models (LLMs) to accelerate the process of implementing parts of software systems. When considering its potential use for ADAS or AD systems in the automotive context, there is a need to systematically assess this new setup: LLMs entail a well-documented set of risks for safety-related systems' development due to their stochastic nature. To reduce the effort for code reviewers to evaluate LLM-generated code, we propose an evaluation pipeline to conduct sanity-checks on the generated code. We compare the performance of six state-of-the-art LLMs (CodeLlama, CodeGemma, DeepSeek-r1, DeepSeek-Coders, Mistral, and GPT-4) on four safety-related programming tasks. Additionally, we qualitatively analyse the most frequent faults generated by these LLMs, creating a failure-mode catalogue to support human reviewers. Finally, the limitations and capabilities of LLMs in code generation, and the use of the proposed pipeline in the existing process, are discussed.
title Large Language Models in Code Co-generation for Safe Autonomous Vehicles
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2505.19658