Unveiling Memorization in Code Models
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Zhou, Zhao, Zhipeng, Wang, Chenyu, Shi, Jieke, Kim, Dongsun, Han, DongGyun, Lo, David |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Finding Safety Violations of AI-Enabled Control Systems through the Lens of Synthesized Proxy Programs
by: Shi, Jieke, et al.
Published: (2024)
by: Shi, Jieke, et al.
Published: (2024)
Gotcha! This Model Uses My Code! Evaluating Membership Leakage Risks in Code Models
by: Yang, Zhou, et al.
Published: (2023)
by: Yang, Zhou, et al.
Published: (2023)
Bridging Expert Knowledge with Deep Learning Techniques for Just-In-Time Defect Prediction
by: Zhou, Xin, et al.
Published: (2024)
by: Zhou, Xin, et al.
Published: (2024)
Multi-LLM Collaboration + Data-Centric Innovation = 2x Better Vulnerability Repair
by: Zhou, Xin, et al.
Published: (2024)
by: Zhou, Xin, et al.
Published: (2024)
PTM4Tag+: Tag Recommendation of Stack Overflow Posts with Pre-trained Models
by: He, Junda, et al.
Published: (2024)
by: He, Junda, et al.
Published: (2024)
Ecosystem of Large Language Models for Code
by: Yang, Zhou, et al.
Published: (2024)
by: Yang, Zhou, et al.
Published: (2024)
Efficient and Green Large Language Models for Software Engineering: Literature Review, Vision, and the Road Ahead
by: Shi, Jieke, et al.
Published: (2024)
by: Shi, Jieke, et al.
Published: (2024)
ACECode: A Reinforcement Learning Framework for Aligning Code Efficiency and Correctness in Code Language Models
by: Yang, Chengran, et al.
Published: (2024)
by: Yang, Chengran, et al.
Published: (2024)
PatchZero: Zero-Shot Automatic Patch Correctness Assessment
by: Zhou, Xin, et al.
Published: (2023)
by: Zhou, Xin, et al.
Published: (2023)
Greening Large Language Models of Code
by: Shi, Jieke, et al.
Published: (2023)
by: Shi, Jieke, et al.
Published: (2023)
Duplicate Bug Report Detection: How Far Are We?
by: Zhang, Ting, et al.
Published: (2022)
by: Zhang, Ting, et al.
Published: (2022)
Think Like Human Developers: Harnessing Community Knowledge for Structured Code Reasoning
by: Yang, Chengran, et al.
Published: (2025)
by: Yang, Chengran, et al.
Published: (2025)
Synthesizing Efficient and Permissive Programmatic Runtime Shields for Neural Policies
by: Shi, Jieke, et al.
Published: (2024)
by: Shi, Jieke, et al.
Published: (2024)
Compiling Code LLMs into Lightweight Executables
by: Shi, Jieke, et al.
Published: (2026)
by: Shi, Jieke, et al.
Published: (2026)
Can LLMs Deobfuscate Binary Code? A Systematic Analysis of Large Language Models into Pseudocode Deobfuscation
by: Hu, Li, et al.
Published: (2026)
by: Hu, Li, et al.
Published: (2026)
Finding Memory Leaks in C/C++ Programs via Neuro-Symbolic Augmented Static Analysis
by: Huang, Huihui, et al.
Published: (2026)
by: Huang, Huihui, et al.
Published: (2026)
Backdoors in Code Summarizers: How Bad Is It?
by: Wang, Chenyu, et al.
Published: (2025)
by: Wang, Chenyu, et al.
Published: (2025)
Curiosity-Driven Testing for Sequential Decision-Making Process
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
Hotfixing Large Language Models for Code
by: Yang, Zhou, et al.
Published: (2024)
by: Yang, Zhou, et al.
Published: (2024)
"My productivity is boosted, but ..." Demystifying Users' Perception on AI Coding Assistants
by: Lyu, Yunbo, et al.
Published: (2025)
by: Lyu, Yunbo, et al.
Published: (2025)
SLICEMATE: Accurate and Scalable Static Program Slicing via LLM-Powered Agents
by: Chang, Jianming, et al.
Published: (2025)
by: Chang, Jianming, et al.
Published: (2025)
From Code to Courtroom: LLMs as the New Software Judges
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
Scrub It Out! Erasing Sensitive Memorization in Code Language Models via Machine Unlearning
by: Chu, Zhaoyang, et al.
Published: (2025)
by: Chu, Zhaoyang, et al.
Published: (2025)
Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
Finding Missing Input Validation in TEEs via LLM-Assisted Symbolic Execution
by: Ma, Chengyan, et al.
Published: (2026)
by: Ma, Chengyan, et al.
Published: (2026)
Learned or Memorized ? Quantifying Memorization Advantage in Code LLMs
by: Euraste, Djiré Albérick, et al.
Published: (2026)
by: Euraste, Djiré Albérick, et al.
Published: (2026)
ESG Reporting Lifecycle Management with Large Language Models and AI Agents
by: Hoang, Thong, et al.
Published: (2026)
by: Hoang, Thong, et al.
Published: (2026)
Hallucination by Code Generation LLMs: Taxonomy, Benchmarks, Mitigation, and Challenges
by: Lee, Yunseo, et al.
Published: (2025)
by: Lee, Yunseo, et al.
Published: (2025)
Automated Repair of TEE Partitioning Issues via DSL-Guided and LLM-Assisted Patching
by: Ma, Chengyan, et al.
Published: (2026)
by: Ma, Chengyan, et al.
Published: (2026)
What You Trust Is Insecure: Demystifying How Developers (Mis)Use Trusted Execution Environments in Practice
by: Niu, Yuqing, et al.
Published: (2025)
by: Niu, Yuqing, et al.
Published: (2025)
Mut4All: Fuzzing Compilers via LLM-Synthesized Mutators Learned from Bug Reports
by: Wang, Bo, et al.
Published: (2025)
by: Wang, Bo, et al.
Published: (2025)
Robustness, Security, Privacy, Explainability, Efficiency, and Usability of Large Language Models for Code
by: Yang, Zhou, et al.
Published: (2024)
by: Yang, Zhou, et al.
Published: (2024)
Adversarial Attacks on Code Models with Discriminative Graph Patterns
by: Nguyen, Thanh-Dat, et al.
Published: (2023)
by: Nguyen, Thanh-Dat, et al.
Published: (2023)
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
AgentSZZ: Teaching the LLM Agent to Play Detective with Bug-Inducing Commits
by: Lyu, Yunbo, et al.
Published: (2026)
by: Lyu, Yunbo, et al.
Published: (2026)
PenForge: On-the-Fly Expert Agent Construction for Automated Penetration Testing
by: Huang, Huihui, et al.
Published: (2026)
by: Huang, Huihui, et al.
Published: (2026)
On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of Code
by: Weyssow, Martin, et al.
Published: (2023)
by: Weyssow, Martin, et al.
Published: (2023)
Hidden Licensing Risks in the LLMware Ecosystem
by: Wang, Bo, et al.
Published: (2026)
by: Wang, Bo, et al.
Published: (2026)
Token Sugar: Making Source Code Sweeter for LLMs through Token-Efficient Shorthand
by: Sun, Zhensu, et al.
Published: (2025)
by: Sun, Zhensu, et al.
Published: (2025)
Bias Unveiled: Investigating Social Bias in LLM-Generated Code
by: Ling, Lin, et al.
Published: (2024)
by: Ling, Lin, et al.
Published: (2024)
Similar Items
-
Finding Safety Violations of AI-Enabled Control Systems through the Lens of Synthesized Proxy Programs
by: Shi, Jieke, et al.
Published: (2024) -
Gotcha! This Model Uses My Code! Evaluating Membership Leakage Risks in Code Models
by: Yang, Zhou, et al.
Published: (2023) -
Bridging Expert Knowledge with Deep Learning Techniques for Just-In-Time Defect Prediction
by: Zhou, Xin, et al.
Published: (2024) -
Multi-LLM Collaboration + Data-Centric Innovation = 2x Better Vulnerability Repair
by: Zhou, Xin, et al.
Published: (2024) -
PTM4Tag+: Tag Recommendation of Stack Overflow Posts with Pre-trained Models
by: He, Junda, et al.
Published: (2024)