Scrub It Out! Erasing Sensitive Memorization in Code Language Models via Machine Unlearning
Fuente:
arXiv
Saved in:
| Main Authors: | Chu, Zhaoyang, Wan, Yao, Zhang, Zhikun, Wang, Di, Yang, Zhou, Zhang, Hongyu, Zhou, Pan, Shi, Xuanhua, Jin, Hai, Lo, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Sensitive Information Leakage in LLMs4Code through Machine Unlearning
by: Gu, Shanzhi, et al.
Published: (2025)
by: Gu, Shanzhi, et al.
Published: (2025)
Graph Neural Networks for Vulnerability Detection: A Counterfactual Explanation
by: Chu, Zhaoyang, et al.
Published: (2024)
by: Chu, Zhaoyang, et al.
Published: (2024)
Gotcha! This Model Uses My Code! Evaluating Membership Leakage Risks in Code Models
by: Yang, Zhou, et al.
Published: (2023)
by: Yang, Zhou, et al.
Published: (2023)
Backdoors in Code Summarizers: How Bad Is It?
by: Wang, Chenyu, et al.
Published: (2025)
by: Wang, Chenyu, et al.
Published: (2025)
Decoding Secret Memorization in Code LLMs Through Token-Level Characterization
by: Nie, Yuqing, et al.
Published: (2024)
by: Nie, Yuqing, et al.
Published: (2024)
Defending Code Language Models against Backdoor Attacks with Deceptive Cross-Entropy Loss
by: Yang, Guang, et al.
Published: (2024)
by: Yang, Guang, et al.
Published: (2024)
Automated TEE Adaptation with LLMs: Identifying, Transforming, and Porting Sensitive Functions in Programs
by: Han, Ruidong, et al.
Published: (2025)
by: Han, Ruidong, et al.
Published: (2025)
Finding Memory Leaks in C/C++ Programs via Neuro-Symbolic Augmented Static Analysis
by: Huang, Huihui, et al.
Published: (2026)
by: Huang, Huihui, et al.
Published: (2026)
Out of Distribution, Out of Luck: How Well Can LLMs Trained on Vulnerability Datasets Detect Top 25 CWE Weaknesses?
by: Li, Yikun, et al.
Published: (2025)
by: Li, Yikun, et al.
Published: (2025)
How Agentic AI Coding Assistants Become the Attacker's Shell
by: Liu, Yue, et al.
Published: (2026)
by: Liu, Yue, et al.
Published: (2026)
Similar but Patched Code Considered Harmful -- The Impact of Similar but Patched Code on Recurring Vulnerability Detection and How to Remove Them
by: Tan, Zixuan, et al.
Published: (2024)
by: Tan, Zixuan, et al.
Published: (2024)
"Your AI, My Shell": Demystifying Prompt Injection Attacks on Agentic AI Coding Editors
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
CEBin: A Cost-Effective Framework for Large-Scale Binary Code Similarity Detection
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
RepoMark: A Data-Usage Auditing Framework for Code Large Language Models
by: Qu, Wenjie, et al.
Published: (2025)
by: Qu, Wenjie, et al.
Published: (2025)
FDI: Attack Neural Code Generation Systems through User Feedback Channel
by: Sun, Zhensu, et al.
Published: (2024)
by: Sun, Zhensu, et al.
Published: (2024)
SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
by: Hu, Qi, et al.
Published: (2026)
by: Hu, Qi, et al.
Published: (2026)
RESCUE: Retrieval Augmented Secure Code Generation
by: Shi, Jiahao, et al.
Published: (2025)
by: Shi, Jiahao, et al.
Published: (2025)
KEENHash: Hashing Programs into Function-Aware Embeddings for Large-Scale Binary Code Similarity Analysis
by: Liu, Zhijie, et al.
Published: (2025)
by: Liu, Zhijie, et al.
Published: (2025)
LLM-Assisted Model-Based Fuzzing of Protocol Implementations
by: Huang, Changze, et al.
Published: (2025)
by: Huang, Changze, et al.
Published: (2025)
An Empirical Study on Remote Code Execution in Machine Learning Model Hosting Ecosystems
by: Siddiq, Mohammed Latif, et al.
Published: (2026)
by: Siddiq, Mohammed Latif, et al.
Published: (2026)
Semantics-Aligned, Curriculum-Driven, and Reasoning-Enhanced Vulnerability Repair Framework
by: Yang, Chengran, et al.
Published: (2025)
by: Yang, Chengran, et al.
Published: (2025)
KVerus: Scalable and Resilient Formal Verification Proof Generation for Rust Code
by: Liu, Yuwei, et al.
Published: (2026)
by: Liu, Yuwei, et al.
Published: (2026)
PPT4J: Patch Presence Test for Java Binaries
by: Pan, Zhiyuan, et al.
Published: (2023)
by: Pan, Zhiyuan, et al.
Published: (2023)
Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks
by: Qu, Yubin, et al.
Published: (2026)
by: Qu, Yubin, et al.
Published: (2026)
SandCell: Sandboxing Rust Beyond Unsafe Code
by: Zhang, Jialun, et al.
Published: (2025)
by: Zhang, Jialun, et al.
Published: (2025)
Concerned with Data Contamination? Assessing Countermeasures in Code Language Model
by: Cao, Jialun, et al.
Published: (2024)
by: Cao, Jialun, et al.
Published: (2024)
Beyond Function-Level Analysis: Context-Aware Reasoning for Inter-Procedural Vulnerability Detection
by: Li, Yikun, et al.
Published: (2026)
by: Li, Yikun, et al.
Published: (2026)
Transferable Backdoor Attacks for Code Models via Sharpness-Aware Adversarial Perturbation
by: Chang, Shuyu, et al.
Published: (2026)
by: Chang, Shuyu, et al.
Published: (2026)
Demonstration Attack against In-Context Learning for Code Intelligence
by: Ge, Yifei, et al.
Published: (2024)
by: Ge, Yifei, et al.
Published: (2024)
CKGFuzzer: LLM-Based Fuzz Driver Generation Enhanced By Code Knowledge Graph
by: Xu, Hanxiang, et al.
Published: (2024)
by: Xu, Hanxiang, et al.
Published: (2024)
DITING: A Static Analyzer for Identifying Bad Partitioning Issues in TEE Applications
by: Ma, Chengyan, et al.
Published: (2025)
by: Ma, Chengyan, et al.
Published: (2025)
MCGMark: An Encodable and Robust Online Watermark for Tracing LLM-Generated Malicious Code
by: Ning, Kaiwen, et al.
Published: (2024)
by: Ning, Kaiwen, et al.
Published: (2024)
D-LiFT: Improving LLM-based Decompiler Backend via Code Quality-driven Fine-tuning
by: Zou, Muqi, et al.
Published: (2025)
by: Zou, Muqi, et al.
Published: (2025)
Models Are Codes: Towards Measuring Malicious Code Poisoning Attacks on Pre-trained Model Hubs
by: Zhao, Jian, et al.
Published: (2024)
by: Zhao, Jian, et al.
Published: (2024)
{A New Hope}: Contextual Privacy Policies for Mobile Applications and An Approach Toward Automated Generation
by: Pan, Shidong, et al.
Published: (2024)
by: Pan, Shidong, et al.
Published: (2024)
CleanVul: Automatic Function-Level Vulnerability Detection in Code Commits Using LLM Heuristics
by: Li, Yikun, et al.
Published: (2024)
by: Li, Yikun, et al.
Published: (2024)
Finding Missing Input Validation in TEEs via LLM-Assisted Symbolic Execution
by: Ma, Chengyan, et al.
Published: (2026)
by: Ma, Chengyan, et al.
Published: (2026)
VERCATION: Precise Vulnerable Open-source Software Version Identification based on Static Analysis and LLM
by: Cheng, Yiran, et al.
Published: (2024)
by: Cheng, Yiran, et al.
Published: (2024)
FidelityGPT: Correcting Decompilation Distortions with Retrieval Augmented Generation
by: Zhou, Zhiping, et al.
Published: (2025)
by: Zhou, Zhiping, et al.
Published: (2025)
DeCoMa: Detecting and Purifying Code Dataset Watermarks through Dual Channel Code Abstraction
by: Xiao, Yuan, et al.
Published: (2025)
by: Xiao, Yuan, et al.
Published: (2025)
Similar Items
-
Mitigating Sensitive Information Leakage in LLMs4Code through Machine Unlearning
by: Gu, Shanzhi, et al.
Published: (2025) -
Graph Neural Networks for Vulnerability Detection: A Counterfactual Explanation
by: Chu, Zhaoyang, et al.
Published: (2024) -
Gotcha! This Model Uses My Code! Evaluating Membership Leakage Risks in Code Models
by: Yang, Zhou, et al.
Published: (2023) -
Backdoors in Code Summarizers: How Bad Is It?
by: Wang, Chenyu, et al.
Published: (2025) -
Decoding Secret Memorization in Code LLMs Through Token-Level Characterization
by: Nie, Yuqing, et al.
Published: (2024)