Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Meifang, Yang, Zhe, Nianchen, Huang, Huang, Yizhan, Li, Yichen, Li, Zihan, Lyu, Michael R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Your Code Secret Belongs to Me: Neural Code Completion Tools Can Memorize Hard-Coded Credentials
by: Huang, Yizhan, et al.
Published: (2023)
by: Huang, Yizhan, et al.
Published: (2023)
Understanding Data Reconstruction Leakage in Federated Learning from a Theoretical Perspective
by: Wang, Zifan, et al.
Published: (2024)
by: Wang, Zifan, et al.
Published: (2024)
Can You Keep a Secret? Involuntary Information Leakage in Language Model Writing
by: Holtzman, Ari, et al.
Published: (2026)
by: Holtzman, Ari, et al.
Published: (2026)
From Data Leak to Secret Misses: The Impact of Data Leakage on Secret Detection Models
by: Soltaniani, Farnaz, et al.
Published: (2026)
by: Soltaniani, Farnaz, et al.
Published: (2026)
Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
by: Lu, Yu-An, et al.
Published: (2026)
by: Lu, Yu-An, et al.
Published: (2026)
Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors
by: Li, Zi, et al.
Published: (2026)
by: Li, Zi, et al.
Published: (2026)
Separating Secrets from Placeholders: A Hybrid CNN-CodeBERT Framework for Three-Class Credential Leakage Detection
by: Baby, Maksuda Bilkis, et al.
Published: (2026)
by: Baby, Maksuda Bilkis, et al.
Published: (2026)
Data Compressibility Quantifies LLM Memorization
by: Huang, Yizhan, et al.
Published: (2025)
by: Huang, Yizhan, et al.
Published: (2025)
Mitigating Sensitive Information Leakage in LLMs4Code through Machine Unlearning
by: Gu, Shanzhi, et al.
Published: (2025)
by: Gu, Shanzhi, et al.
Published: (2025)
ShadowCode: Towards (Automatic) External Prompt Injection Attack against Code LLMs
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
Protect Your Secrets: Understanding and Measuring Data Exposure in VSCode Extensions
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
by: Tan, Yixin, et al.
Published: (2025)
by: Tan, Yixin, et al.
Published: (2025)
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
by: Chen, Qirui, et al.
Published: (2026)
by: Chen, Qirui, et al.
Published: (2026)
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
by: Li, Yige, et al.
Published: (2026)
by: Li, Yige, et al.
Published: (2026)
Learning to Generate Secure Code via Token-Level Rewards
by: Quan, Jiazheng, et al.
Published: (2026)
by: Quan, Jiazheng, et al.
Published: (2026)
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
by: Cao, Bochuan, et al.
Published: (2025)
by: Cao, Bochuan, et al.
Published: (2025)
A First Look At Efficient And Secure On-Device LLM Inference Against KV Leakage
by: Yang, Huan, et al.
Published: (2024)
by: Yang, Huan, et al.
Published: (2024)
Credential Leakage in LLM Agent Skills: A Large-Scale Empirical Study
by: Chen, Zhihao, et al.
Published: (2026)
by: Chen, Zhihao, et al.
Published: (2026)
EquaCode: A Multi-Strategy Jailbreak Approach for Large Language Models via Equation Solving and Code Completion
by: Liang, Zhen, et al.
Published: (2025)
by: Liang, Zhen, et al.
Published: (2025)
CompLeak: Deep Learning Model Compression Exacerbates Privacy Leakage
by: Li, Na, et al.
Published: (2025)
by: Li, Na, et al.
Published: (2025)
SoK: On Gradient Leakage in Federated Learning
by: Du, Jiacheng, et al.
Published: (2024)
by: Du, Jiacheng, et al.
Published: (2024)
Beyond TVLA: Anderson-Darling Leakage Assessment for Neural Network Side-Channel Leakage Detection
by: Mikulec, Ján, et al.
Published: (2026)
by: Mikulec, Ján, et al.
Published: (2026)
LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights
by: Sheng, Ze, et al.
Published: (2025)
by: Sheng, Ze, et al.
Published: (2025)
TokenMark: A Modality-Agnostic Watermark for Pre-trained Transformers
by: Xu, Hengyuan, et al.
Published: (2024)
by: Xu, Hengyuan, et al.
Published: (2024)
Argus: A Multi-Agent Sensitive Information Leakage Detection Framework Based on Hierarchical Reference Relationships
by: Wang, Bin, et al.
Published: (2025)
by: Wang, Bin, et al.
Published: (2025)
Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data
by: ElZemity, Adel, et al.
Published: (2025)
by: ElZemity, Adel, et al.
Published: (2025)
SoK: Understanding (New) Security Issues Across AI4Code Use Cases
by: Wu, Qilong, et al.
Published: (2025)
by: Wu, Qilong, et al.
Published: (2025)
How Does Naming Affect LLMs on Code Analysis Tasks?
by: Wang, Zhilong, et al.
Published: (2023)
by: Wang, Zhilong, et al.
Published: (2023)
Do Coding Agents Understand Least-Privilege Authorization?
by: Yan, Zheng, et al.
Published: (2026)
by: Yan, Zheng, et al.
Published: (2026)
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
by: Wang, Liwen, et al.
Published: (2025)
by: Wang, Liwen, et al.
Published: (2025)
Shadows in the Code: Exploring the Risks and Defenses of LLM-based Multi-Agent Software Development Systems
by: Wang, Xiaoqing, et al.
Published: (2025)
by: Wang, Xiaoqing, et al.
Published: (2025)
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
by: Chen, Wenyu, et al.
Published: (2026)
by: Chen, Wenyu, et al.
Published: (2026)
Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment
by: Xie, Yuchong, et al.
Published: (2025)
by: Xie, Yuchong, et al.
Published: (2025)
Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs
by: Liu, Jinbo, et al.
Published: (2025)
by: Liu, Jinbo, et al.
Published: (2025)
Quantifying Association Capabilities of Large Language Models and Its Implications on Privacy Leakage
by: Shao, Hanyin, et al.
Published: (2023)
by: Shao, Hanyin, et al.
Published: (2023)
Membership Inference Attacks on Tokenizers of Large Language Models
by: Tong, Meng, et al.
Published: (2025)
by: Tong, Meng, et al.
Published: (2025)
Emerging Cyber Attack Risks of Medical AI Agents
by: Qiu, Jianing, et al.
Published: (2025)
by: Qiu, Jianing, et al.
Published: (2025)
Understanding Privacy Risks in Code Models Through Training Dynamics: A Causal Approach
by: Yang, Hua, et al.
Published: (2025)
by: Yang, Hua, et al.
Published: (2025)
Cracking IoT Security: Can LLMs Outsmart Static Analysis Tools?
by: Quantrill, Jason, et al.
Published: (2026)
by: Quantrill, Jason, et al.
Published: (2026)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
by: Jiang, Peihai, et al.
Published: (2025)
by: Jiang, Peihai, et al.
Published: (2025)
Similar Items
-
Your Code Secret Belongs to Me: Neural Code Completion Tools Can Memorize Hard-Coded Credentials
by: Huang, Yizhan, et al.
Published: (2023) -
Understanding Data Reconstruction Leakage in Federated Learning from a Theoretical Perspective
by: Wang, Zifan, et al.
Published: (2024) -
Can You Keep a Secret? Involuntary Information Leakage in Language Model Writing
by: Holtzman, Ari, et al.
Published: (2026) -
From Data Leak to Secret Misses: The Impact of Data Leakage on Secret Detection Models
by: Soltaniani, Farnaz, et al.
Published: (2026) -
Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
by: Lu, Yu-An, et al.
Published: (2026)