A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lian, Keke, Wang, Bin, Zhang, Lei, Chen, Libo, Wang, Junjie, Zhao, Ziming, Yang, Yujiu, Lin, Miaoqian, Duan, Haotong, Zhao, Haoran, Liao, Shuang, Guo, Mingda, Quan, Jiazheng, Zhong, Yilu, He, Chenhao, Chen, Zichuan, Wu, Jie, Li, Haoling, Li, Zhaoxuan, Yu, Jiongchi, Li, Hui, Zhang, Dong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916956090662912
author Lian, Keke
Wang, Bin
Zhang, Lei
Chen, Libo
Wang, Junjie
Zhao, Ziming
Yang, Yujiu
Lin, Miaoqian
Duan, Haotong
Zhao, Haoran
Liao, Shuang
Guo, Mingda
Quan, Jiazheng
Zhong, Yilu
He, Chenhao
Chen, Zichuan
Wu, Jie
Li, Haoling
Li, Zhaoxuan
Yu, Jiongchi
Li, Hui
Zhang, Dong
author_facet Lian, Keke
Wang, Bin
Zhang, Lei
Chen, Libo
Wang, Junjie
Zhao, Ziming
Yang, Yujiu
Lin, Miaoqian
Duan, Haotong
Zhao, Haoran
Liao, Shuang
Guo, Mingda
Quan, Jiazheng
Zhong, Yilu
He, Chenhao
Chen, Zichuan
Wu, Jie
Li, Haoling
Li, Zhaoxuan
Yu, Jiongchi
Li, Hui
Zhang, Dong
contents The increasing adoption of large language models (LLMs) in software engineering necessitates rigorous security evaluation of their generated code. However, existing benchmarks often lack relevance to real-world AI-assisted programming scenarios, making them inadequate for assessing the practical security risks associated with AI-generated code in production environments. To address this gap, we introduce A.S.E (AI Code Generation Security Evaluation), a repository-level evaluation benchmark designed to closely mirror real-world AI programming tasks, offering a comprehensive and reliable framework for assessing the security of AI-generated code. Our evaluation of leading LLMs on A.S.E reveals several key findings. In particular, current LLMs still struggle with secure coding. The complexity in repository-level scenarios presents challenges for LLMs that typically perform well on snippet-level tasks. Moreover, a larger reasoning budget does not necessarily lead to better code generation. These observations offer valuable insights into the current state of AI code generation and help developers identify the most suitable models for practical tasks. They also lay the groundwork for refining LLMs to generate secure and efficient code in real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18106
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
Lian, Keke
Wang, Bin
Zhang, Lei
Chen, Libo
Wang, Junjie
Zhao, Ziming
Yang, Yujiu
Lin, Miaoqian
Duan, Haotong
Zhao, Haoran
Liao, Shuang
Guo, Mingda
Quan, Jiazheng
Zhong, Yilu
He, Chenhao
Chen, Zichuan
Wu, Jie
Li, Haoling
Li, Zhaoxuan
Yu, Jiongchi
Li, Hui
Zhang, Dong
Software Engineering
Artificial Intelligence
The increasing adoption of large language models (LLMs) in software engineering necessitates rigorous security evaluation of their generated code. However, existing benchmarks often lack relevance to real-world AI-assisted programming scenarios, making them inadequate for assessing the practical security risks associated with AI-generated code in production environments. To address this gap, we introduce A.S.E (AI Code Generation Security Evaluation), a repository-level evaluation benchmark designed to closely mirror real-world AI programming tasks, offering a comprehensive and reliable framework for assessing the security of AI-generated code. Our evaluation of leading LLMs on A.S.E reveals several key findings. In particular, current LLMs still struggle with secure coding. The complexity in repository-level scenarios presents challenges for LLMs that typically perform well on snippet-level tasks. Moreover, a larger reasoning budget does not necessarily lead to better code generation. These observations offer valuable insights into the current state of AI code generation and help developers identify the most suitable models for practical tasks. They also lay the groundwork for refining LLMs to generate secure and efficient code in real-world applications.
title A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2508.18106