Reconstruction of Personally Identifiable Information from Supervised Finetuned Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Furukawa, Sae, Oprea, Alina |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TMI! Finetuned Models Leak Private Information from their Pretraining Data
por: Abascal, John, et al.
Publicado: (2023)
por: Abascal, John, et al.
Publicado: (2023)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
por: Naseh, Ali, et al.
Publicado: (2025)
por: Naseh, Ali, et al.
Publicado: (2025)
User Inference Attacks on Large Language Models
por: Kandpal, Nikhil, et al.
Publicado: (2023)
por: Kandpal, Nikhil, et al.
Publicado: (2023)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
por: Chaudhari, Harsh, et al.
Publicado: (2024)
por: Chaudhari, Harsh, et al.
Publicado: (2024)
Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security
por: Naseh, Ali, et al.
Publicado: (2025)
por: Naseh, Ali, et al.
Publicado: (2025)
Randomized Masked Finetuning: An Efficient Way to Mitigate Memorization of PIIs in LLMs
por: Joshi, Kunj, et al.
Publicado: (2025)
por: Joshi, Kunj, et al.
Publicado: (2025)
Identifying Models Behind Text-to-Image Leaderboards
por: Naseh, Ali, et al.
Publicado: (2026)
por: Naseh, Ali, et al.
Publicado: (2026)
Adversarial Inception Backdoor Attacks against Reinforcement Learning
por: Rathbun, Ethan, et al.
Publicado: (2024)
por: Rathbun, Ethan, et al.
Publicado: (2024)
SleeperNets: Universal Backdoor Poisoning Attacks Against Reinforcement Learning Agents
por: Rathbun, Ethan, et al.
Publicado: (2024)
por: Rathbun, Ethan, et al.
Publicado: (2024)
Personal Information Parroting in Language Models
por: Subramani, Nishant, et al.
Publicado: (2026)
por: Subramani, Nishant, et al.
Publicado: (2026)
Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation
por: Naseh, Ali, et al.
Publicado: (2025)
por: Naseh, Ali, et al.
Publicado: (2025)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
por: Halawi, Danny, et al.
Publicado: (2024)
por: Halawi, Danny, et al.
Publicado: (2024)
Cascading Adversarial Bias from Injection to Distillation in Language Models
por: Chaudhari, Harsh, et al.
Publicado: (2025)
por: Chaudhari, Harsh, et al.
Publicado: (2025)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
por: Jiang, Weisen, et al.
Publicado: (2025)
por: Jiang, Weisen, et al.
Publicado: (2025)
When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents
por: Xu, Xiaoyu, et al.
Publicado: (2026)
por: Xu, Xiaoyu, et al.
Publicado: (2026)
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
por: Suri, Anshuman, et al.
Publicado: (2025)
por: Suri, Anshuman, et al.
Publicado: (2025)
Synthesizing Tight Privacy and Accuracy Bounds via Weighted Model Counting
por: Oakley, Lisa, et al.
Publicado: (2024)
por: Oakley, Lisa, et al.
Publicado: (2024)
Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors
por: Sun, Luze, et al.
Publicado: (2026)
por: Sun, Luze, et al.
Publicado: (2026)
Automated CVE Analysis: Harnessing Machine Learning In Designing Question-Answering Models For Cybersecurity Information Extraction
por: Faruk, Tanjim Bin
Publicado: (2024)
por: Faruk, Tanjim Bin
Publicado: (2024)
Protecting Copyrighted Material with Unique Identifiers in Large Language Model Training
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
Detecting Pretraining Data from Large Language Models
por: Shi, Weijia, et al.
Publicado: (2023)
por: Shi, Weijia, et al.
Publicado: (2023)
Backdoor Attacks in Peer-to-Peer Federated Learning
por: Syros, Georgios, et al.
Publicado: (2023)
por: Syros, Georgios, et al.
Publicado: (2023)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
por: Suri, Omar Farooq Khan, et al.
Publicado: (2025)
por: Suri, Omar Farooq Khan, et al.
Publicado: (2025)
Beware Untrusted Simulators -- Reward-Free Backdoor Attacks in Reinforcement Learning
por: Rathbun, Ethan, et al.
Publicado: (2026)
por: Rathbun, Ethan, et al.
Publicado: (2026)
Privately Learning from Graphs with Applications in Fine-tuning Large Language Models
por: Yin, Haoteng, et al.
Publicado: (2024)
por: Yin, Haoteng, et al.
Publicado: (2024)
On the Effectiveness of Membership Inference in Targeted Data Extraction from Large Language Models
por: Sahili, Ali Al, et al.
Publicado: (2025)
por: Sahili, Ali Al, et al.
Publicado: (2025)
Watermarking Language Models through Language Models
por: Dasgupta, Agnibh, et al.
Publicado: (2024)
por: Dasgupta, Agnibh, et al.
Publicado: (2024)
Model Provenance Testing for Large Language Models
por: Nikolic, Ivica, et al.
Publicado: (2025)
por: Nikolic, Ivica, et al.
Publicado: (2025)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
por: Kaneko, Masahiro, et al.
Publicado: (2025)
por: Kaneko, Masahiro, et al.
Publicado: (2025)
Malware Classification from Memory Dumps Using Machine Learning, Transformers, and Large Language Models
por: Dweib, Areej, et al.
Publicado: (2025)
por: Dweib, Areej, et al.
Publicado: (2025)
What Does the Server See? Understanding Privacy Leakage from Large Language Models in Split Inference
por: Fan, Mingyuan, et al.
Publicado: (2026)
por: Fan, Mingyuan, et al.
Publicado: (2026)
Black-Box Privacy Attacks on Shared Representations in Multitask Learning
por: Abascal, John, et al.
Publicado: (2025)
por: Abascal, John, et al.
Publicado: (2025)
UTrace: Poisoning Forensics for Private Collaborative Learning
por: Rose, Evan, et al.
Publicado: (2024)
por: Rose, Evan, et al.
Publicado: (2024)
Quantitative Resilience Modeling for Autonomous Cyber Defense
por: Cadet, Xavier, et al.
Publicado: (2025)
por: Cadet, Xavier, et al.
Publicado: (2025)
Towards the Anonymization of the Language Modeling
por: Boutet, Antoine, et al.
Publicado: (2025)
por: Boutet, Antoine, et al.
Publicado: (2025)
On the Learnability of Watermarks for Language Models
por: Gu, Chenchen, et al.
Publicado: (2023)
por: Gu, Chenchen, et al.
Publicado: (2023)
A Watermark for Large Language Models
por: Kirchenbauer, John, et al.
Publicado: (2023)
por: Kirchenbauer, John, et al.
Publicado: (2023)
Localizing Paragraph Memorization in Language Models
por: Stoehr, Niklas, et al.
Publicado: (2024)
por: Stoehr, Niklas, et al.
Publicado: (2024)
Are PPO-ed Language Models Hackable?
por: Anand, Suraj, et al.
Publicado: (2024)
por: Anand, Suraj, et al.
Publicado: (2024)
On the Reliability of Watermarks for Large Language Models
por: Kirchenbauer, John, et al.
Publicado: (2023)
por: Kirchenbauer, John, et al.
Publicado: (2023)
Ejemplares similares
-
TMI! Finetuned Models Leak Private Information from their Pretraining Data
por: Abascal, John, et al.
Publicado: (2023) -
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
por: Naseh, Ali, et al.
Publicado: (2025) -
User Inference Attacks on Large Language Models
por: Kandpal, Nikhil, et al.
Publicado: (2023) -
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
por: Chaudhari, Harsh, et al.
Publicado: (2024) -
Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security
por: Naseh, Ali, et al.
Publicado: (2025)