Extracting Memorized Training Data via Decomposition
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Ellen, Vellore, Anu, Chang, Amy, Mura, Raffaele, Nelson, Blaine, Kassianik, Paul, Karbasi, Amin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report
by: Kassianik, Paul, et al.
Published: (2025)
by: Kassianik, Paul, et al.
Published: (2025)
Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
by: Weerawardhena, Sajana, et al.
Published: (2025)
by: Weerawardhena, Sajana, et al.
Published: (2025)
Llama-3.1-FoundationAI-SecurityLLM-Reasoning-8B Technical Report
by: Yang, Zhuoran, et al.
Published: (2026)
by: Yang, Zhuoran, et al.
Published: (2026)
A Framework for Rapidly Developing and Deploying Protection Against Large Language Model Attacks
by: Swanda, Adam, et al.
Published: (2025)
by: Swanda, Adam, et al.
Published: (2025)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
Extracting Training Data from Diffusion Language Models via Infilling
by: Wang, Yihan, et al.
Published: (2026)
by: Wang, Yihan, et al.
Published: (2026)
Unveiling Privacy, Memorization, and Input Curvature Links
by: Ravikumar, Deepak, et al.
Published: (2024)
by: Ravikumar, Deepak, et al.
Published: (2024)
Position: Privacy Is Not Just Memorization!
by: Mireshghallah, Niloofar, et al.
Published: (2025)
by: Mireshghallah, Niloofar, et al.
Published: (2025)
PureGen: Universal Data Purification for Train-Time Poison Defense via Generative Model Dynamics
by: Bhat, Sunay, et al.
Published: (2024)
by: Bhat, Sunay, et al.
Published: (2024)
Probe-Geometry Alignment: Erasing the Cross-Sequence Memorization Signature Below Chance
by: Rupa, Anamika Paul, et al.
Published: (2026)
by: Rupa, Anamika Paul, et al.
Published: (2026)
Uncertainty-Aware Hardware Trojan Detection Using Multimodal Deep Learning
by: Vishwakarma, Rahul, et al.
Published: (2024)
by: Vishwakarma, Rahul, et al.
Published: (2024)
Private Memorization Editing: Turning Memorization into a Defense to Strengthen Data Privacy in Large Language Models
by: Ruzzetti, Elena Sofia, et al.
Published: (2025)
by: Ruzzetti, Elena Sofia, et al.
Published: (2025)
Reconstructing Training Data from Adapter-based Federated Large Language Models
by: Chen, Silong, et al.
Published: (2026)
by: Chen, Silong, et al.
Published: (2026)
Invisible Backdoor Attack Through Singular Value Decomposition
by: Chen, Wenmin, et al.
Published: (2024)
by: Chen, Wenmin, et al.
Published: (2024)
Unlocking Memorization in Large Language Models with Dynamic Soft Prompting
by: Wang, Zhepeng, et al.
Published: (2024)
by: Wang, Zhepeng, et al.
Published: (2024)
Pandora's White-Box: Precise Training Data Detection and Extraction in Large Language Models
by: Wang, Jeffrey G., et al.
Published: (2024)
by: Wang, Jeffrey G., et al.
Published: (2024)
Evidencing Unauthorized Training Data from AI Generated Content using Information Isotopes
by: Tao, Qi, et al.
Published: (2025)
by: Tao, Qi, et al.
Published: (2025)
Multi-class Classifier based Failure Prediction with Artificial and Anonymous Training for Data Privacy
by: Das, Dibakar, et al.
Published: (2022)
by: Das, Dibakar, et al.
Published: (2022)
What's Pulling the Strings? Evaluating Integrity and Attribution in AI Training and Inference through Concept Shift
by: Chang, Jiamin, et al.
Published: (2025)
by: Chang, Jiamin, et al.
Published: (2025)
Improving Clean Accuracy via a Tangent-Space Perspective on Adversarial Training
by: Yi, Bongsoo, et al.
Published: (2024)
by: Yi, Bongsoo, et al.
Published: (2024)
Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models
by: Sivashanmugam, Sathesh P.
Published: (2025)
by: Sivashanmugam, Sathesh P.
Published: (2025)
Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models
by: Dong, Yihong, et al.
Published: (2024)
by: Dong, Yihong, et al.
Published: (2024)
TEN-GUARD: Tensor Decomposition for Backdoor Attack Detection in Deep Neural Networks
by: Hossain, Khondoker Murad, et al.
Published: (2024)
by: Hossain, Khondoker Murad, et al.
Published: (2024)
CLASP: Training-Free LLM-Assisted Source Code Watermarking via Semantic-Preserving Transformations
by: Xu, Rui, et al.
Published: (2025)
by: Xu, Rui, et al.
Published: (2025)
Training on Fake Labels: Mitigating Label Leakage in Split Learning via Secure Dimension Transformation
by: Jiang, Yukun, et al.
Published: (2024)
by: Jiang, Yukun, et al.
Published: (2024)
Detecting Training Data of Large Language Models via Expectation Maximization
by: Kim, Gyuwan, et al.
Published: (2024)
by: Kim, Gyuwan, et al.
Published: (2024)
Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection
by: Zeng, Qirun, et al.
Published: (2025)
by: Zeng, Qirun, et al.
Published: (2025)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
by: Banayeeanzade, Amin, et al.
Published: (2026)
by: Banayeeanzade, Amin, et al.
Published: (2026)
Data Unlearning Beyond Uniform Forgetting via Diffusion Time and Frequency Selection
by: Park, Jinseong, et al.
Published: (2025)
by: Park, Jinseong, et al.
Published: (2025)
Data-Free Privacy-Preserving for LLMs via Model Inversion and Selective Unlearning
by: Zhou, Xinjie, et al.
Published: (2026)
by: Zhou, Xinjie, et al.
Published: (2026)
Adaptive Discounting of Training Time Attacks
by: Bector, Ridhima, et al.
Published: (2024)
by: Bector, Ridhima, et al.
Published: (2024)
KnowGraph: Knowledge-Enabled Anomaly Detection via Logical Reasoning on Graph Data
by: Zhou, Andy, et al.
Published: (2024)
by: Zhou, Andy, et al.
Published: (2024)
ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite Thinking
by: Li, Yunzhe, et al.
Published: (2025)
by: Li, Yunzhe, et al.
Published: (2025)
Optimistic Verifiable Training by Controlling Hardware Nondeterminism
by: Srivastava, Megha, et al.
Published: (2024)
by: Srivastava, Megha, et al.
Published: (2024)
Leveraging RAG for Training-Free Alignment of LLMs
by: Halloran, John T.
Published: (2026)
by: Halloran, John T.
Published: (2026)
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks
by: Halloran, John T., et al.
Published: (2026)
by: Halloran, John T., et al.
Published: (2026)
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
by: Chen, Shuhao, et al.
Published: (2026)
by: Chen, Shuhao, et al.
Published: (2026)
JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
by: Chu, Junjie, et al.
Published: (2025)
by: Chu, Junjie, et al.
Published: (2025)
Information Theoretic Adversarial Training of Large Language Models
by: Zhang, Yiwei, et al.
Published: (2026)
by: Zhang, Yiwei, et al.
Published: (2026)
Similar Items
-
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023) -
Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report
by: Kassianik, Paul, et al.
Published: (2025) -
Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
by: Weerawardhena, Sajana, et al.
Published: (2025) -
Llama-3.1-FoundationAI-SecurityLLM-Reasoning-8B Technical Report
by: Yang, Zhuoran, et al.
Published: (2026) -
A Framework for Rapidly Developing and Deploying Protection Against Large Language Model Attacks
by: Swanda, Adam, et al.
Published: (2025)