Can You Really Trust Code Copilots? Evaluating Large Language Models from a Code Security Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mou, Yutao, Deng, Xiao, Luo, Yuxiao, Zhang, Shikun, Ye, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913839973400576
author Mou, Yutao
Deng, Xiao
Luo, Yuxiao
Zhang, Shikun
Ye, Wei
author_facet Mou, Yutao
Deng, Xiao
Luo, Yuxiao
Zhang, Shikun
Ye, Wei
contents Code security and usability are both essential for various coding assistant applications driven by large language models (LLMs). Current code security benchmarks focus solely on single evaluation task and paradigm, such as code completion and generation, lacking comprehensive assessment across dimensions like secure code generation, vulnerability repair and discrimination. In this paper, we first propose CoV-Eval, a multi-task benchmark covering various tasks such as code completion, vulnerability repair, vulnerability detection and classification, for comprehensive evaluation of LLM code security. Besides, we developed VC-Judge, an improved judgment model that aligns closely with human experts and can review LLM-generated programs for vulnerabilities in a more efficient and reliable way. We conduct a comprehensive evaluation of 20 proprietary and open-source LLMs. Overall, while most LLMs identify vulnerable codes well, they still tend to generate insecure codes and struggle with recognizing specific vulnerability types and performing repairs. Extensive experiments and qualitative analyses reveal key challenges and optimization directions, offering insights for future research in LLM code security.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10494
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can You Really Trust Code Copilots? Evaluating Large Language Models from a Code Security Perspective
Mou, Yutao
Deng, Xiao
Luo, Yuxiao
Zhang, Shikun
Ye, Wei
Computation and Language
Code security and usability are both essential for various coding assistant applications driven by large language models (LLMs). Current code security benchmarks focus solely on single evaluation task and paradigm, such as code completion and generation, lacking comprehensive assessment across dimensions like secure code generation, vulnerability repair and discrimination. In this paper, we first propose CoV-Eval, a multi-task benchmark covering various tasks such as code completion, vulnerability repair, vulnerability detection and classification, for comprehensive evaluation of LLM code security. Besides, we developed VC-Judge, an improved judgment model that aligns closely with human experts and can review LLM-generated programs for vulnerabilities in a more efficient and reliable way. We conduct a comprehensive evaluation of 20 proprietary and open-source LLMs. Overall, while most LLMs identify vulnerable codes well, they still tend to generate insecure codes and struggle with recognizing specific vulnerability types and performing repairs. Extensive experiments and qualitative analyses reveal key challenges and optimization directions, offering insights for future research in LLM code security.
title Can You Really Trust Code Copilots? Evaluating Large Language Models from a Code Security Perspective
topic Computation and Language
url https://arxiv.org/abs/2505.10494