Train in Vain: Functionality-Preserving Poisoning to Prevent Unauthorized Use of Code Datasets
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Yuan, Wang, Jiaming, Chen, Yuchen, Song, Wei, Sun, Jun, Ma, Shiqing, Mu, Yanzhou, Zhai, Juan, Fang, Chunrong, Dong, Jin Song, Chen, Zhenyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeCoMa: Detecting and Purifying Code Dataset Watermarks through Dual Channel Code Abstraction
by: Xiao, Yuan, et al.
Published: (2025)
by: Xiao, Yuan, et al.
Published: (2025)
DuCodeMark: Dual-Purpose Code Dataset Watermarking via Style-Aware Watermark-Poison Design
by: Chen, Yuchen, et al.
Published: (2026)
by: Chen, Yuchen, et al.
Published: (2026)
Probing Privacy Leaks in LLM-based Code Generation via Test Generation
by: Ge, Yifei, et al.
Published: (2026)
by: Ge, Yifei, et al.
Published: (2026)
Security of Language Models for Code: A Systematic Literature Review
by: Chen, Yuchen, et al.
Published: (2024)
by: Chen, Yuchen, et al.
Published: (2024)
Show Me Your Code! Kill Code Poisoning: A Lightweight Method Based on Code Naturalness
by: Sun, Weisong, et al.
Published: (2025)
by: Sun, Weisong, et al.
Published: (2025)
Demonstration Attack against In-Context Learning for Code Intelligence
by: Ge, Yifei, et al.
Published: (2024)
by: Ge, Yifei, et al.
Published: (2024)
Exploring the Security Threats of Knowledge Base Poisoning in Retrieval-Augmented Code Generation
by: Lin, Bo, et al.
Published: (2025)
by: Lin, Bo, et al.
Published: (2025)
Models Are Codes: Towards Measuring Malicious Code Poisoning Attacks on Pre-trained Model Hubs
by: Zhao, Jian, et al.
Published: (2024)
by: Zhao, Jian, et al.
Published: (2024)
Detecting Stealthy Data Poisoning Attacks in AI Code Generators
by: Improta, Cristina
Published: (2025)
by: Improta, Cristina
Published: (2025)
When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
by: Peng, Yibo, et al.
Published: (2025)
by: Peng, Yibo, et al.
Published: (2025)
SCAFFOLD-CEGIS: Preventing Latent Security Degradation in LLM-Driven Iterative Code Refinement
by: Chen, Yi, et al.
Published: (2026)
by: Chen, Yi, et al.
Published: (2026)
The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code Generation
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
Poisoning Programs by Un-Repairing Code: Security Concerns of AI-generated Code
by: Improta, Cristina
Published: (2024)
by: Improta, Cristina
Published: (2024)
Eliminating Backdoors in Neural Code Models for Secure Code Understanding
by: Sun, Weisong, et al.
Published: (2024)
by: Sun, Weisong, et al.
Published: (2024)
Towards Robust Detection of Open Source Software Supply Chain Poisoning Attacks in Industry Environments
by: Zheng, Xinyi, et al.
Published: (2024)
by: Zheng, Xinyi, et al.
Published: (2024)
Towards Efficient Verification of Constant-Time Cryptographic Implementations
by: Cai, Luwei, et al.
Published: (2024)
by: Cai, Luwei, et al.
Published: (2024)
Hybrid Privacy Policy-Code Consistency Check using Knowledge Graphs and LLMs
by: Mao, Zhenyu, et al.
Published: (2025)
by: Mao, Zhenyu, et al.
Published: (2025)
FDI: Attack Neural Code Generation Systems through User Feedback Channel
by: Sun, Zhensu, et al.
Published: (2024)
by: Sun, Zhensu, et al.
Published: (2024)
Enhancing Pre-Trained Language Models for Vulnerability Detection via Semantic-Preserving Data Augmentation
by: Qi, Weiliang, et al.
Published: (2024)
by: Qi, Weiliang, et al.
Published: (2024)
MegaVul: A C/C++ Vulnerability Dataset with Comprehensive Code Representation
by: Ni, Chao, et al.
Published: (2024)
by: Ni, Chao, et al.
Published: (2024)
CodableLLM: Automating Decompiled and Source Code Mapping for LLM Dataset Generation
by: Manuel, Dylan, et al.
Published: (2025)
by: Manuel, Dylan, et al.
Published: (2025)
XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants
by: Štorek, Adam, et al.
Published: (2025)
by: Štorek, Adam, et al.
Published: (2025)
CEBin: A Cost-Effective Framework for Large-Scale Binary Code Similarity Detection
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models
by: Jiang, Weipeng, et al.
Published: (2026)
by: Jiang, Weipeng, et al.
Published: (2026)
Detecting Data Poisoning in Code Generation LLMs via Black-Box, Vulnerability-Oriented Scanning
by: Yan, Shenao, et al.
Published: (2026)
by: Yan, Shenao, et al.
Published: (2026)
I Know Who Clones Your Code: Interpretable Smart Contract Similarity Detection
by: Liu, Zhenguang, et al.
Published: (2025)
by: Liu, Zhenguang, et al.
Published: (2025)
Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning
by: Huang, Charoes, et al.
Published: (2026)
by: Huang, Charoes, et al.
Published: (2026)
REBENCH: A Procedural, Fair-by-Construction Benchmark for LLMs on Stripped-Binary Types and Names (Extended Version)
by: Won, Jun Yeon, et al.
Published: (2026)
by: Won, Jun Yeon, et al.
Published: (2026)
Out of Distribution, Out of Luck: How Well Can LLMs Trained on Vulnerability Datasets Detect Top 25 CWE Weaknesses?
by: Li, Yikun, et al.
Published: (2025)
by: Li, Yikun, et al.
Published: (2025)
FCGHunter: Towards Evaluating Robustness of Graph-Based Android Malware Detection
by: Song, Shiwen, et al.
Published: (2025)
by: Song, Shiwen, et al.
Published: (2025)
Semantic Consensus Decoding: Backdoor Defense for Verilog Code Generation
by: Yang, Guang, et al.
Published: (2026)
by: Yang, Guang, et al.
Published: (2026)
Towards Privacy-Preserving Code Generation: Differentially Private Code Language Models
by: Catal, Melih, et al.
Published: (2025)
by: Catal, Melih, et al.
Published: (2025)
Optimal Circuit Synthesis of Linear Codes for Error Detection and Correction
by: Yang, Xi, et al.
Published: (2026)
by: Yang, Xi, et al.
Published: (2026)
Constraint-based Adversarial Example Synthesis
by: Yu, Fang, et al.
Published: (2024)
by: Yu, Fang, et al.
Published: (2024)
Defending Code Language Models against Backdoor Attacks with Deceptive Cross-Entropy Loss
by: Yang, Guang, et al.
Published: (2024)
by: Yang, Guang, et al.
Published: (2024)
Large Language Models for Code Analysis: Do LLMs Really Do Their Job?
by: Fang, Chongzhou, et al.
Published: (2023)
by: Fang, Chongzhou, et al.
Published: (2023)
Understanding Privacy Risks in Code Models Through Training Dynamics: A Causal Approach
by: Yang, Hua, et al.
Published: (2025)
by: Yang, Hua, et al.
Published: (2025)
A Study of Malware Prevention in Linux Distributions
by: Vu, Duc-Ly, et al.
Published: (2024)
by: Vu, Duc-Ly, et al.
Published: (2024)
From Detection to Prevention: Explaining Security-Critical Code to Avoid Vulnerabilities
by: Krishnamurthy, Ranjith, et al.
Published: (2026)
by: Krishnamurthy, Ranjith, et al.
Published: (2026)
KEENHash: Hashing Programs into Function-Aware Embeddings for Large-Scale Binary Code Similarity Analysis
by: Liu, Zhijie, et al.
Published: (2025)
by: Liu, Zhijie, et al.
Published: (2025)
Similar Items
-
DeCoMa: Detecting and Purifying Code Dataset Watermarks through Dual Channel Code Abstraction
by: Xiao, Yuan, et al.
Published: (2025) -
DuCodeMark: Dual-Purpose Code Dataset Watermarking via Style-Aware Watermark-Poison Design
by: Chen, Yuchen, et al.
Published: (2026) -
Probing Privacy Leaks in LLM-based Code Generation via Test Generation
by: Ge, Yifei, et al.
Published: (2026) -
Security of Language Models for Code: A Systematic Literature Review
by: Chen, Yuchen, et al.
Published: (2024) -
Show Me Your Code! Kill Code Poisoning: A Lightweight Method Based on Code Naturalness
by: Sun, Weisong, et al.
Published: (2025)