Detecting Data Poisoning in Code Generation LLMs via Black-Box, Vulnerability-Oriented Scanning
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Shenao, Ahmed, Shimaa, Jin, Shan, Arora, Sunpreet S., Cai, Yiwei, Wang, Yizhen, Hong, Yuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
by: Yan, Shenao, et al.
Published: (2024)
by: Yan, Shenao, et al.
Published: (2024)
Models Are Codes: Towards Measuring Malicious Code Poisoning Attacks on Pre-trained Model Hubs
by: Zhao, Jian, et al.
Published: (2024)
by: Zhao, Jian, et al.
Published: (2024)
Towards Robust Detection of Open Source Software Supply Chain Poisoning Attacks in Industry Environments
by: Zheng, Xinyi, et al.
Published: (2024)
by: Zheng, Xinyi, et al.
Published: (2024)
Detecting Stealthy Data Poisoning Attacks in AI Code Generators
by: Improta, Cristina
Published: (2025)
by: Improta, Cristina
Published: (2025)
Harnessing the Power of LLMs in Source Code Vulnerability Detection
by: Mahyari, Andrew A
Published: (2024)
by: Mahyari, Andrew A
Published: (2024)
CleanVul: Automatic Function-Level Vulnerability Detection in Code Commits Using LLM Heuristics
by: Li, Yikun, et al.
Published: (2024)
by: Li, Yikun, et al.
Published: (2024)
Train in Vain: Functionality-Preserving Poisoning to Prevent Unauthorized Use of Code Datasets
by: Xiao, Yuan, et al.
Published: (2026)
by: Xiao, Yuan, et al.
Published: (2026)
Impact of Code Transformation on Detection of Smart Contract Vulnerabilities
by: Manh, Cuong Tran, et al.
Published: (2024)
by: Manh, Cuong Tran, et al.
Published: (2024)
Unsupervised Binary Code Translation with Application to Code Similarity Detection and Vulnerability Discovery
by: Ahmad, Iftakhar, et al.
Published: (2024)
by: Ahmad, Iftakhar, et al.
Published: (2024)
How Code Representation Shapes False-Positive Dynamics in Cross-Language LLM Vulnerability Detection
by: Chen, Maofei, et al.
Published: (2026)
by: Chen, Maofei, et al.
Published: (2026)
Lightweight Vulnerability Detection from Code Metrics and Token Features
by: Chiu, Chun Yin
Published: (2026)
by: Chiu, Chun Yin
Published: (2026)
StagedVulBERT: Multi-Granular Vulnerability Detection with a Novel Pre-trained Code Model
by: Jiang, Yuan, et al.
Published: (2024)
by: Jiang, Yuan, et al.
Published: (2024)
Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning
by: Huang, Charoes, et al.
Published: (2026)
by: Huang, Charoes, et al.
Published: (2026)
A Slicing-Based Approach for Detecting and Patching Vulnerable Code Clones
by: Alomari, Hakam, et al.
Published: (2025)
by: Alomari, Hakam, et al.
Published: (2025)
Does the Vulnerability Threaten Our Projects? Automated Vulnerable API Detection for Third-Party Libraries
by: Zhang, Fangyuan, et al.
Published: (2024)
by: Zhang, Fangyuan, et al.
Published: (2024)
Detecting Protracted Vulnerabilities in Open Source Projects
by: Sridharkumar, Arjun, et al.
Published: (2026)
by: Sridharkumar, Arjun, et al.
Published: (2026)
AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications
by: Yang, Ruozhao, et al.
Published: (2026)
by: Yang, Ruozhao, et al.
Published: (2026)
When Labels Are Scarce: A Systematic Mapping of Label-Efficient Code Vulnerability Detection
by: Khalal, Noor, et al.
Published: (2026)
by: Khalal, Noor, et al.
Published: (2026)
Vulnerability Detection with Interprocedural Context in Multiple Languages: Assessing Effectiveness and Cost of Modern LLMs
by: Lira, Kevin, et al.
Published: (2026)
by: Lira, Kevin, et al.
Published: (2026)
Similar but Patched Code Considered Harmful -- The Impact of Similar but Patched Code on Recurring Vulnerability Detection and How to Remove Them
by: Tan, Zixuan, et al.
Published: (2024)
by: Tan, Zixuan, et al.
Published: (2024)
MegaVul: A C/C++ Vulnerability Dataset with Comprehensive Code Representation
by: Ni, Chao, et al.
Published: (2024)
by: Ni, Chao, et al.
Published: (2024)
Exploring the Security Threats of Knowledge Base Poisoning in Retrieval-Augmented Code Generation
by: Lin, Bo, et al.
Published: (2025)
by: Lin, Bo, et al.
Published: (2025)
VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications
by: Zhu, Hao, et al.
Published: (2025)
by: Zhu, Hao, et al.
Published: (2025)
"Elementary, My Dear Watson." Detecting Malicious Skills via Neuro-Symbolic Reasoning across Heterogeneous Artifacts
by: Wang, Shenao, et al.
Published: (2026)
by: Wang, Shenao, et al.
Published: (2026)
Static Security Vulnerability Scanning of Proprietary and Open-Source Software: An Adaptable Process with Variants and Results
by: Cusick, James J.
Published: (2025)
by: Cusick, James J.
Published: (2025)
HALURust: Exploiting Hallucinations of Large Language Models to Detect Vulnerabilities in Rust
by: Luo, Yu, et al.
Published: (2025)
by: Luo, Yu, et al.
Published: (2025)
Uncover the Premeditated Attacks: Detecting Exploitable Reentrancy Vulnerabilities by Identifying Attacker Contracts
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
MiniScope: Automated UI Exploration and Privacy Inconsistency Detection of MiniApps via Two-phase Iterative Hybrid Analysis
by: Wang, Shenao, et al.
Published: (2024)
by: Wang, Shenao, et al.
Published: (2024)
Is GitHub's Copilot as Bad as Humans at Introducing Vulnerabilities in Code?
by: Asare, Owura, et al.
Published: (2022)
by: Asare, Owura, et al.
Published: (2022)
Guiding Symbolic Execution with Static Analysis and LLMs for Vulnerability Discovery
by: Shafiuzzaman, Md, et al.
Published: (2026)
by: Shafiuzzaman, Md, et al.
Published: (2026)
WACANA: A Concolic Analyzer for Detecting On-chain Data Vulnerabilities in WASM Smart Contracts
by: Wang, Wansen, et al.
Published: (2024)
by: Wang, Wansen, et al.
Published: (2024)
Game Rewards Vulnerabilities: Software Vulnerability Detection with Zero-Sum Game and Prototype Learning
by: Wen, Xin-Cheng, et al.
Published: (2024)
by: Wen, Xin-Cheng, et al.
Published: (2024)
Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs
by: Sternfeld, Alexander, et al.
Published: (2026)
by: Sternfeld, Alexander, et al.
Published: (2026)
TPSQLi: Test Prioritization for SQL Injection Vulnerability Detection in Web Applications
by: Yang, Guan-Yan, et al.
Published: (2025)
by: Yang, Guan-Yan, et al.
Published: (2025)
Tricking Retrievers with Influential Tokens: An Efficient Black-Box Corpus Poisoning Attack
by: Wang, Cheng, et al.
Published: (2025)
by: Wang, Cheng, et al.
Published: (2025)
On the Effectiveness of Function-Level Vulnerability Detectors for Inter-Procedural Vulnerabilities
by: Li, Zhen, et al.
Published: (2024)
by: Li, Zhen, et al.
Published: (2024)
Out of Distribution, Out of Luck: How Well Can LLMs Trained on Vulnerability Datasets Detect Top 25 CWE Weaknesses?
by: Li, Yikun, et al.
Published: (2025)
by: Li, Yikun, et al.
Published: (2025)
Do Fine-Tuned LLMs Understand Vulnerabilities? An Investigation into the Semantic Trap
by: Huang, Feiyang, et al.
Published: (2026)
by: Huang, Feiyang, et al.
Published: (2026)
Vital: Vulnerability-Oriented Symbolic Execution via Type-Unsafe Pointer-Guided Monte Carlo Tree Search
by: Tu, Haoxin, et al.
Published: (2024)
by: Tu, Haoxin, et al.
Published: (2024)
AutoVulnPHP: LLM-Powered Two-Stage PHP Vulnerability Detection and Automated Localization
by: Wang, Zhiqiang, et al.
Published: (2026)
by: Wang, Zhiqiang, et al.
Published: (2026)
Similar Items
-
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
by: Yan, Shenao, et al.
Published: (2024) -
Models Are Codes: Towards Measuring Malicious Code Poisoning Attacks on Pre-trained Model Hubs
by: Zhao, Jian, et al.
Published: (2024) -
Towards Robust Detection of Open Source Software Supply Chain Poisoning Attacks in Industry Environments
by: Zheng, Xinyi, et al.
Published: (2024) -
Detecting Stealthy Data Poisoning Attacks in AI Code Generators
by: Improta, Cristina
Published: (2025) -
Harnessing the Power of LLMs in Source Code Vulnerability Detection
by: Mahyari, Andrew A
Published: (2024)