Randomized Masked Finetuning: An Efficient Way to Mitigate Memorization of PIIs in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Joshi, Kunj, Smith, David A. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
Localizing Paragraph Memorization in Language Models
di: Stoehr, Niklas, et al.
Pubblicazione: (2024)
di: Stoehr, Niklas, et al.
Pubblicazione: (2024)
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
di: Mathew, Yohan, et al.
Pubblicazione: (2024)
di: Mathew, Yohan, et al.
Pubblicazione: (2024)
Data-centric NLP Backdoor Defense from the Lens of Memorization
di: Wang, Zhenting, et al.
Pubblicazione: (2024)
di: Wang, Zhenting, et al.
Pubblicazione: (2024)
Reconstruction of Personally Identifiable Information from Supervised Finetuned Models
di: Furukawa, Sae, et al.
Pubblicazione: (2026)
di: Furukawa, Sae, et al.
Pubblicazione: (2026)
Watermarking Needs Input Repetition Masking
di: Khachaturov, David, et al.
Pubblicazione: (2025)
di: Khachaturov, David, et al.
Pubblicazione: (2025)
Leaner Training, Lower Leakage: Revisiting Memorization in LLM Fine-Tuning with LoRA
di: Wang, Fei, et al.
Pubblicazione: (2025)
di: Wang, Fei, et al.
Pubblicazione: (2025)
Position: Privacy Is Not Just Memorization!
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2025)
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2025)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
di: Betley, Jan, et al.
Pubblicazione: (2025)
di: Betley, Jan, et al.
Pubblicazione: (2025)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
di: Halawi, Danny, et al.
Pubblicazione: (2024)
di: Halawi, Danny, et al.
Pubblicazione: (2024)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
di: Assogba, Yannick, et al.
Pubblicazione: (2026)
di: Assogba, Yannick, et al.
Pubblicazione: (2026)
Unlocking Memorization in Large Language Models with Dynamic Soft Prompting
di: Wang, Zhepeng, et al.
Pubblicazione: (2024)
di: Wang, Zhepeng, et al.
Pubblicazione: (2024)
Sanitize Your Responses: Mitigating Privacy Leakage in Large Language Models
di: Fu, Wenjie, et al.
Pubblicazione: (2025)
di: Fu, Wenjie, et al.
Pubblicazione: (2025)
How Vulnerable Are Edge LLMs?
di: Ding, Ao, et al.
Pubblicazione: (2026)
di: Ding, Ao, et al.
Pubblicazione: (2026)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
di: Jiang, Weisen, et al.
Pubblicazione: (2025)
di: Jiang, Weisen, et al.
Pubblicazione: (2025)
UCD: Unlearning in LLMs via Contrastive Decoding
di: Suriyakumar, Vinith M., et al.
Pubblicazione: (2025)
di: Suriyakumar, Vinith M., et al.
Pubblicazione: (2025)
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
di: Dotsinski, Asen, et al.
Pubblicazione: (2026)
di: Dotsinski, Asen, et al.
Pubblicazione: (2026)
Coercing LLMs to do and reveal (almost) anything
di: Geiping, Jonas, et al.
Pubblicazione: (2024)
di: Geiping, Jonas, et al.
Pubblicazione: (2024)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
di: Aldahoul, Nouar, et al.
Pubblicazione: (2025)
di: Aldahoul, Nouar, et al.
Pubblicazione: (2025)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
di: Sivapiromrat, Sanhanat, et al.
Pubblicazione: (2025)
di: Sivapiromrat, Sanhanat, et al.
Pubblicazione: (2025)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
di: Zhao, Xuandong, et al.
Pubblicazione: (2024)
di: Zhao, Xuandong, et al.
Pubblicazione: (2024)
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
di: Wallace, Eric, et al.
Pubblicazione: (2024)
di: Wallace, Eric, et al.
Pubblicazione: (2024)
AdaptDel: Adaptable Deletion Rate Randomized Smoothing for Certified Robustness
di: Huang, Zhuoqun, et al.
Pubblicazione: (2025)
di: Huang, Zhuoqun, et al.
Pubblicazione: (2025)
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
di: Wang, Linlin, et al.
Pubblicazione: (2025)
di: Wang, Linlin, et al.
Pubblicazione: (2025)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
di: Brown, Hannah, et al.
Pubblicazione: (2024)
di: Brown, Hannah, et al.
Pubblicazione: (2024)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
di: Price, Sara, et al.
Pubblicazione: (2024)
di: Price, Sara, et al.
Pubblicazione: (2024)
Can We Infer Confidential Properties of Training Data from LLMs?
di: Huang, Pengrun, et al.
Pubblicazione: (2025)
di: Huang, Pengrun, et al.
Pubblicazione: (2025)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
di: Suri, Omar Farooq Khan, et al.
Pubblicazione: (2025)
di: Suri, Omar Farooq Khan, et al.
Pubblicazione: (2025)
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
di: Li, Ran, et al.
Pubblicazione: (2025)
di: Li, Ran, et al.
Pubblicazione: (2025)
Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs
di: Atwal, Tevin, et al.
Pubblicazione: (2025)
di: Atwal, Tevin, et al.
Pubblicazione: (2025)
SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
di: Meeus, Matthieu, et al.
Pubblicazione: (2024)
di: Meeus, Matthieu, et al.
Pubblicazione: (2024)
LMO-DP: Optimizing the Randomization Mechanism for Differentially Private Fine-Tuning (Large) Language Models
di: Yang, Qin, et al.
Pubblicazione: (2024)
di: Yang, Qin, et al.
Pubblicazione: (2024)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
di: Yoon, Do-hyeon, et al.
Pubblicazione: (2025)
di: Yoon, Do-hyeon, et al.
Pubblicazione: (2025)
Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content
di: Liu, Bing, et al.
Pubblicazione: (2026)
di: Liu, Bing, et al.
Pubblicazione: (2026)
LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTI
di: Schwartz, Yuval, et al.
Pubblicazione: (2024)
di: Schwartz, Yuval, et al.
Pubblicazione: (2024)
CodeCloak: A Method for Evaluating and Mitigating Code Leakage by LLM Code Assistants
di: Noah, Amit Finkman, et al.
Pubblicazione: (2024)
di: Noah, Amit Finkman, et al.
Pubblicazione: (2024)
PromptScreen: Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline
di: Rao, Akshaj Prashanth, et al.
Pubblicazione: (2025)
di: Rao, Akshaj Prashanth, et al.
Pubblicazione: (2025)
No Free Lunch in LLM Watermarking: Trade-offs in Watermarking Design Choices
di: Pang, Qi, et al.
Pubblicazione: (2024)
di: Pang, Qi, et al.
Pubblicazione: (2024)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
di: Xiong, Alexander, et al.
Pubblicazione: (2025) -
Localizing Paragraph Memorization in Language Models
di: Stoehr, Niklas, et al.
Pubblicazione: (2024) -
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
di: Mathew, Yohan, et al.
Pubblicazione: (2024) -
Data-centric NLP Backdoor Defense from the Lens of Memorization
di: Wang, Zhenting, et al.
Pubblicazione: (2024) -
Reconstruction of Personally Identifiable Information from Supervised Finetuned Models
di: Furukawa, Sae, et al.
Pubblicazione: (2026)