DeCoMa: Detecting and Purifying Code Dataset Watermarks through Dual Channel Code Abstraction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Yuan, Chen, Yuchen, Ma, Shiqing, Huang, Haocheng, Fang, Chunrong, Chen, Yanwei, Sun, Weisong, Zhu, Yunfeng, Zhang, Xiaofang, Chen, Zhenyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PuzzleMark: Implicit Jigsaw Learning for Robust Code Dataset Watermarking in Neural Code Completion Models
von: Huang, Haocheng, et al.
Veröffentlicht: (2026)
von: Huang, Haocheng, et al.
Veröffentlicht: (2026)
Probing Privacy Leaks in LLM-based Code Generation via Test Generation
von: Ge, Yifei, et al.
Veröffentlicht: (2026)
von: Ge, Yifei, et al.
Veröffentlicht: (2026)
DuCodeMark: Dual-Purpose Code Dataset Watermarking via Style-Aware Watermark-Poison Design
von: Chen, Yuchen, et al.
Veröffentlicht: (2026)
von: Chen, Yuchen, et al.
Veröffentlicht: (2026)
Security of Language Models for Code: A Systematic Literature Review
von: Chen, Yuchen, et al.
Veröffentlicht: (2024)
von: Chen, Yuchen, et al.
Veröffentlicht: (2024)
Train in Vain: Functionality-Preserving Poisoning to Prevent Unauthorized Use of Code Datasets
von: Xiao, Yuan, et al.
Veröffentlicht: (2026)
von: Xiao, Yuan, et al.
Veröffentlicht: (2026)
Demonstration Attack against In-Context Learning for Code Intelligence
von: Ge, Yifei, et al.
Veröffentlicht: (2024)
von: Ge, Yifei, et al.
Veröffentlicht: (2024)
Show Me Your Code! Kill Code Poisoning: A Lightweight Method Based on Code Naturalness
von: Sun, Weisong, et al.
Veröffentlicht: (2025)
von: Sun, Weisong, et al.
Veröffentlicht: (2025)
Eliminating Backdoors in Neural Code Models for Secure Code Understanding
von: Sun, Weisong, et al.
Veröffentlicht: (2024)
von: Sun, Weisong, et al.
Veröffentlicht: (2024)
Enhancing and Reporting Robustness Boundary of Neural Code Models for Intelligent Code Understanding
von: Han, Tingxu, et al.
Veröffentlicht: (2026)
von: Han, Tingxu, et al.
Veröffentlicht: (2026)
MCGMark: An Encodable and Robust Online Watermark for Tracing LLM-Generated Malicious Code
von: Ning, Kaiwen, et al.
Veröffentlicht: (2024)
von: Ning, Kaiwen, et al.
Veröffentlicht: (2024)
ESALE: Enhancing Code-Summary Alignment Learning for Source Code Summarization
von: Fang, Chunrong, et al.
Veröffentlicht: (2024)
von: Fang, Chunrong, et al.
Veröffentlicht: (2024)
FDI: Attack Neural Code Generation Systems through User Feedback Channel
von: Sun, Zhensu, et al.
Veröffentlicht: (2024)
von: Sun, Zhensu, et al.
Veröffentlicht: (2024)
Test Script Intention Generation for Mobile Application via GUI Image and Code Understanding
von: Yu, Shengcheng, et al.
Veröffentlicht: (2021)
von: Yu, Shengcheng, et al.
Veröffentlicht: (2021)
TransformCode: A Contrastive Learning Framework for Code Embedding via Subtree Transformation
von: Xian, Zixiang, et al.
Veröffentlicht: (2023)
von: Xian, Zixiang, et al.
Veröffentlicht: (2023)
No Man is an Island: Towards Fully Automatic Programming by Code Search, Code Generation and Program Repair
von: Zhang, Quanjun, et al.
Veröffentlicht: (2024)
von: Zhang, Quanjun, et al.
Veröffentlicht: (2024)
Models Are Codes: Towards Measuring Malicious Code Poisoning Attacks on Pre-trained Model Hubs
von: Zhao, Jian, et al.
Veröffentlicht: (2024)
von: Zhao, Jian, et al.
Veröffentlicht: (2024)
Hybrid Privacy Policy-Code Consistency Check using Knowledge Graphs and LLMs
von: Mao, Zhenyu, et al.
Veröffentlicht: (2025)
von: Mao, Zhenyu, et al.
Veröffentlicht: (2025)
Semantic Consensus Decoding: Backdoor Defense for Verilog Code Generation
von: Yang, Guang, et al.
Veröffentlicht: (2026)
von: Yang, Guang, et al.
Veröffentlicht: (2026)
Defending Code Language Models against Backdoor Attacks with Deceptive Cross-Entropy Loss
von: Yang, Guang, et al.
Veröffentlicht: (2024)
von: Yang, Guang, et al.
Veröffentlicht: (2024)
Towards Code Watermarking with Dual-Channel Transformations
von: Yang, Borui, et al.
Veröffentlicht: (2023)
von: Yang, Borui, et al.
Veröffentlicht: (2023)
Tightening Robustness Verification of MaxPool-based Neural Networks via Minimizing the Over-Approximation Zone
von: Xiao, Yuan, et al.
Veröffentlicht: (2022)
von: Xiao, Yuan, et al.
Veröffentlicht: (2022)
Exploring the Security Threats of Knowledge Base Poisoning in Retrieval-Augmented Code Generation
von: Lin, Bo, et al.
Veröffentlicht: (2025)
von: Lin, Bo, et al.
Veröffentlicht: (2025)
MegaVul: A C/C++ Vulnerability Dataset with Comprehensive Code Representation
von: Ni, Chao, et al.
Veröffentlicht: (2024)
von: Ni, Chao, et al.
Veröffentlicht: (2024)
CodableLLM: Automating Decompiled and Source Code Mapping for LLM Dataset Generation
von: Manuel, Dylan, et al.
Veröffentlicht: (2025)
von: Manuel, Dylan, et al.
Veröffentlicht: (2025)
The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code Generation
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2025)
CEBin: A Cost-Effective Framework for Large-Scale Binary Code Similarity Detection
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
SCAFFOLD-CEGIS: Preventing Latent Security Degradation in LLM-Driven Iterative Code Refinement
von: Chen, Yi, et al.
Veröffentlicht: (2026)
von: Chen, Yi, et al.
Veröffentlicht: (2026)
WildCode: An Empirical Analysis of Code Generated by ChatGPT
von: Khanmohammadi, Kobra, et al.
Veröffentlicht: (2025)
von: Khanmohammadi, Kobra, et al.
Veröffentlicht: (2025)
An Effective Approach to Embedding Source Code by Combining Large Language and Sentence Embedding Models
von: Xian, Zixiang, et al.
Veröffentlicht: (2024)
von: Xian, Zixiang, et al.
Veröffentlicht: (2024)
Towards Automated Crowdsourced Testing via Personified-LLM
von: Yu, Shengcheng, et al.
Veröffentlicht: (2026)
von: Yu, Shengcheng, et al.
Veröffentlicht: (2026)
Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution
von: Shang, Ye, et al.
Veröffentlicht: (2026)
von: Shang, Ye, et al.
Veröffentlicht: (2026)
An Empirical Study on the Effectiveness of Large Language Models for Binary Code Understanding
von: Shang, Xiuwei, et al.
Veröffentlicht: (2025)
von: Shang, Xiuwei, et al.
Veröffentlicht: (2025)
Security Weaknesses of Copilot-Generated Code in GitHub Projects: An Empirical Study
von: Fu, Yujia, et al.
Veröffentlicht: (2023)
von: Fu, Yujia, et al.
Veröffentlicht: (2023)
How to Compare the Security of Code Written by Humans to LLM-generated Code
von: Balebako, Rebecca, et al.
Veröffentlicht: (2026)
von: Balebako, Rebecca, et al.
Veröffentlicht: (2026)
SecCodePRM: A Process Reward Model for Code Security
von: Yu, Weichen, et al.
Veröffentlicht: (2026)
von: Yu, Weichen, et al.
Veröffentlicht: (2026)
Give LLMs a Security Course: Securing Retrieval-Augmented Code Generation via Knowledge Injection
von: Lin, Bo, et al.
Veröffentlicht: (2025)
von: Lin, Bo, et al.
Veröffentlicht: (2025)
CKGFuzzer: LLM-Based Fuzz Driver Generation Enhanced By Code Knowledge Graph
von: Xu, Hanxiang, et al.
Veröffentlicht: (2024)
von: Xu, Hanxiang, et al.
Veröffentlicht: (2024)
Large Language Models for Code Analysis: Do LLMs Really Do Their Job?
von: Fang, Chongzhou, et al.
Veröffentlicht: (2023)
von: Fang, Chongzhou, et al.
Veröffentlicht: (2023)
How Code Representation Shapes False-Positive Dynamics in Cross-Language LLM Vulnerability Detection
von: Chen, Maofei, et al.
Veröffentlicht: (2026)
von: Chen, Maofei, et al.
Veröffentlicht: (2026)
Unsupervised Binary Code Translation with Application to Code Similarity Detection and Vulnerability Discovery
von: Ahmad, Iftakhar, et al.
Veröffentlicht: (2024)
von: Ahmad, Iftakhar, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PuzzleMark: Implicit Jigsaw Learning for Robust Code Dataset Watermarking in Neural Code Completion Models
von: Huang, Haocheng, et al.
Veröffentlicht: (2026) -
Probing Privacy Leaks in LLM-based Code Generation via Test Generation
von: Ge, Yifei, et al.
Veröffentlicht: (2026) -
DuCodeMark: Dual-Purpose Code Dataset Watermarking via Style-Aware Watermark-Poison Design
von: Chen, Yuchen, et al.
Veröffentlicht: (2026) -
Security of Language Models for Code: A Systematic Literature Review
von: Chen, Yuchen, et al.
Veröffentlicht: (2024) -
Train in Vain: Functionality-Preserving Poisoning to Prevent Unauthorized Use of Code Datasets
von: Xiao, Yuan, et al.
Veröffentlicht: (2026)