Enhancing and Reporting Robustness Boundary of Neural Code Models for Intelligent Code Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Tingxu, Song, Wei, Sun, Weisong, Wu, Hao, Fang, Chunrong, Xiao, Yuan, Zhang, Xiaofang, Chen, Zhenyu, Liu, Yang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911543751344128
author Han, Tingxu
Song, Wei
Sun, Weisong
Wu, Hao
Fang, Chunrong
Xiao, Yuan
Zhang, Xiaofang
Chen, Zhenyu
Liu, Yang
author_facet Han, Tingxu
Song, Wei
Sun, Weisong
Wu, Hao
Fang, Chunrong
Xiao, Yuan
Zhang, Xiaofang
Chen, Zhenyu
Liu, Yang
contents With the development of deep learning, Neural Code Models (NCMs) such as CodeBERT and CodeLlama are widely used for code understanding tasks, including defect detection and code classification. However, recent studies have revealed that NCMs are vulnerable to adversarial examples, inputs with subtle perturbations that induce incorrect predictions while remaining difficult to detect. Existing defenses address this issue via data augmentation to empirically improve robustness, but they are costly, offer no theoretical robustness guarantees, and typically require white-box access to model internals, such as gradients. To address the above challenges, we propose ENBECOME, a novel black-box training-free and lightweight adversarial defense. ENBECOME is designed to both enhance empirical robustness and report certified robustness boundaries for NCMs. ENBECOME operates solely during inference, introducing random, semantics-preserving perturbations to input code snippets to smooth the NCM's decision boundaries. This smoothing enables ENBECOME to formally certify a robustness radius within which adversarial examples can never induce misclassification, a property known as certified robustness. We conduct comprehensive experiments across multiple NCM architectures and tasks. Results show that ENBECOME significantly reduces attack success rates while maintaining high accuracy. For example, in defect detection, it reduces the average ASR from 42.43% to 9.74% with only a 0.29% drop in accuracy. Results show that ENBECOME significantly reduces attack success rates while maintaining high accuracy. For example, in defect detection, it reduces the average ASR from 42.43% to 9.74% with only a 0.29% drop in accuracy. Furthermore, ENBECOME achieves an average certified robustness radius of 1.63, meaning that adversarial modifications to no more than 1.63 identifiers are provably ineffective.
format Preprint
id arxiv_https___arxiv_org_abs_2603_24119
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Enhancing and Reporting Robustness Boundary of Neural Code Models for Intelligent Code Understanding
Han, Tingxu
Song, Wei
Sun, Weisong
Wu, Hao
Fang, Chunrong
Xiao, Yuan
Zhang, Xiaofang
Chen, Zhenyu
Liu, Yang
Software Engineering
With the development of deep learning, Neural Code Models (NCMs) such as CodeBERT and CodeLlama are widely used for code understanding tasks, including defect detection and code classification. However, recent studies have revealed that NCMs are vulnerable to adversarial examples, inputs with subtle perturbations that induce incorrect predictions while remaining difficult to detect. Existing defenses address this issue via data augmentation to empirically improve robustness, but they are costly, offer no theoretical robustness guarantees, and typically require white-box access to model internals, such as gradients. To address the above challenges, we propose ENBECOME, a novel black-box training-free and lightweight adversarial defense. ENBECOME is designed to both enhance empirical robustness and report certified robustness boundaries for NCMs. ENBECOME operates solely during inference, introducing random, semantics-preserving perturbations to input code snippets to smooth the NCM's decision boundaries. This smoothing enables ENBECOME to formally certify a robustness radius within which adversarial examples can never induce misclassification, a property known as certified robustness. We conduct comprehensive experiments across multiple NCM architectures and tasks. Results show that ENBECOME significantly reduces attack success rates while maintaining high accuracy. For example, in defect detection, it reduces the average ASR from 42.43% to 9.74% with only a 0.29% drop in accuracy. Results show that ENBECOME significantly reduces attack success rates while maintaining high accuracy. For example, in defect detection, it reduces the average ASR from 42.43% to 9.74% with only a 0.29% drop in accuracy. Furthermore, ENBECOME achieves an average certified robustness radius of 1.63, meaning that adversarial modifications to no more than 1.63 identifiers are provably ineffective.
title Enhancing and Reporting Robustness Boundary of Neural Code Models for Intelligent Code Understanding
topic Software Engineering
url https://arxiv.org/abs/2603.24119