Discovering Universal Activation Directions for PII Leakage in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Marchyok, Leo, Coalson, Zachary, Keum, Sungho, Son, Sooel, Hong, Sanghyun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Asking Forever: Universal Activations Behind Turn Amplification in Conversational LLMs
by: Coalson, Zachary, et al.
Published: (2026)
by: Coalson, Zachary, et al.
Published: (2026)
Fail-Closed Alignment for Large Language Models
by: Coalson, Zachary, et al.
Published: (2026)
by: Coalson, Zachary, et al.
Published: (2026)
Hard Work Does Not Always Pay Off: Poisoning Attacks on Neural Architecture Search
by: Coalson, Zachary, et al.
Published: (2024)
by: Coalson, Zachary, et al.
Published: (2024)
IF-GUIDE: Influence Function-Guided Detoxification of LLMs
by: Coalson, Zachary, et al.
Published: (2025)
by: Coalson, Zachary, et al.
Published: (2025)
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
by: Wen, Yuxin, et al.
Published: (2024)
by: Wen, Yuxin, et al.
Published: (2024)
PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips
by: Coalson, Zachary, et al.
Published: (2024)
by: Coalson, Zachary, et al.
Published: (2024)
Adaptive PII Mitigation Framework for Large Language Models
by: Asthana, Shubhi, et al.
Published: (2025)
by: Asthana, Shubhi, et al.
Published: (2025)
Modeling Neural Networks with Privacy Using Neural Stochastic Differential Equations
by: Hong, Sanghyun, et al.
Published: (2025)
by: Hong, Sanghyun, et al.
Published: (2025)
Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models
by: Sivashanmugam, Sathesh P.
Published: (2025)
by: Sivashanmugam, Sathesh P.
Published: (2025)
Analysis of Privacy Leakage in Federated Large Language Models
by: Vu, Minh N., et al.
Published: (2024)
by: Vu, Minh N., et al.
Published: (2024)
Information Leakage from Embedding in Large Language Models
by: Wan, Zhipeng, et al.
Published: (2024)
by: Wan, Zhipeng, et al.
Published: (2024)
PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage
by: Nakka, Krishna Kanth, et al.
Published: (2025)
by: Nakka, Krishna Kanth, et al.
Published: (2025)
Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models
by: Zhang, Boyang, et al.
Published: (2025)
by: Zhang, Boyang, et al.
Published: (2025)
Gradient-Free Privacy Leakage in Federated Language Models through Selective Weight Tampering
by: Rashid, Md Rafi Ur, et al.
Published: (2023)
by: Rashid, Md Rafi Ur, et al.
Published: (2023)
Hessian-aware Training for Enhancing DNNs Resilience to Parameter Corruptions
by: Prato, Tahmid Hasan, et al.
Published: (2025)
by: Prato, Tahmid Hasan, et al.
Published: (2025)
Understanding Deep Gradient Leakage via Inversion Influence Functions
by: Zhang, Haobo, et al.
Published: (2023)
by: Zhang, Haobo, et al.
Published: (2023)
MADCAT: Combating Malware Detection Under Concept Drift with Test-Time Adaptation
by: Roh, Eunjin, et al.
Published: (2025)
by: Roh, Eunjin, et al.
Published: (2025)
Real-Time Privacy Risk Measurement with Privacy Tokens for Gradient Leakage
by: Meng, Jiayang, et al.
Published: (2025)
by: Meng, Jiayang, et al.
Published: (2025)
Information Leakage from Data Updates in Machine Learning Models
by: Hui, Tian, et al.
Published: (2023)
by: Hui, Tian, et al.
Published: (2023)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
by: Nakka, Krishna Kanth, et al.
Published: (2024)
by: Nakka, Krishna Kanth, et al.
Published: (2024)
Discovering Spoofing Attempts on Language Model Watermarks
by: Gloaguen, Thibaud, et al.
Published: (2024)
by: Gloaguen, Thibaud, et al.
Published: (2024)
Location Leakage in Federated Signal Maps
by: Bakopoulou, Evita, et al.
Published: (2021)
by: Bakopoulou, Evita, et al.
Published: (2021)
DeepLeak: Privacy Enhancing Hardening of Model Explanations Against Membership Leakage
by: Hmida, Firas Ben, et al.
Published: (2026)
by: Hmida, Firas Ben, et al.
Published: (2026)
Learning to Localize Leakage of Cryptographic Sensitive Variables
by: Gammell, Jimmy, et al.
Published: (2025)
by: Gammell, Jimmy, et al.
Published: (2025)
Sanitize Your Responses: Mitigating Privacy Leakage in Large Language Models
by: Fu, Wenjie, et al.
Published: (2025)
by: Fu, Wenjie, et al.
Published: (2025)
Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage
by: Rashid, Md Rafi Ur, et al.
Published: (2024)
by: Rashid, Md Rafi Ur, et al.
Published: (2024)
Defending Large Language Models Against Attacks With Residual Stream Activation Analysis
by: Kawasaki, Amelia, et al.
Published: (2024)
by: Kawasaki, Amelia, et al.
Published: (2024)
A Survey of What to Share in Federated Learning: Perspectives on Model Utility, Privacy Leakage, and Communication Efficiency
by: Shao, Jiawei, et al.
Published: (2023)
by: Shao, Jiawei, et al.
Published: (2023)
Tracing Privacy Leakage of Language Models to Training Data via Adjusted Influence Functions
by: Liu, Jinxin, et al.
Published: (2024)
by: Liu, Jinxin, et al.
Published: (2024)
PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing
by: Hughes, Anthony, et al.
Published: (2025)
by: Hughes, Anthony, et al.
Published: (2025)
Refiner: Data Refining against Gradient Leakage Attacks in Federated Learning
by: Fan, Mingyuan, et al.
Published: (2022)
by: Fan, Mingyuan, et al.
Published: (2022)
Investigating Privacy Leakage in Dimensionality Reduction Methods via Reconstruction Attack
by: Lumbut, Chayadon, et al.
Published: (2024)
by: Lumbut, Chayadon, et al.
Published: (2024)
ARES: Scalable and Practical Gradient Inversion Attack in Federated Learning through Activation Recovery
by: Gong, Zirui, et al.
Published: (2026)
by: Gong, Zirui, et al.
Published: (2026)
Discovering Command and Control Channels Using Reinforcement Learning
by: Wang, Cheng, et al.
Published: (2024)
by: Wang, Cheng, et al.
Published: (2024)
Leakage Safe Graph Features for Interpretable Fraud Detection in Temporal Transaction Networks
by: Khaleghpour, Hamideh, et al.
Published: (2026)
by: Khaleghpour, Hamideh, et al.
Published: (2026)
Unveiling Client Privacy Leakage from Public Dataset Usage in Federated Distillation
by: Shi, Haonan, et al.
Published: (2025)
by: Shi, Haonan, et al.
Published: (2025)
Functional Encryption in Secure Neural Network Training: Data Leakage and Practical Mitigations
by: Ioniţă, Alexandru, et al.
Published: (2025)
by: Ioniţă, Alexandru, et al.
Published: (2025)
Random Gradient Masking as a Defensive Measure to Deep Leakage in Federated Learning
by: Kim, Joon, et al.
Published: (2024)
by: Kim, Joon, et al.
Published: (2024)
Synth-MIA: A Testbed for Auditing Privacy Leakage in Tabular Data Synthesis
by: Ward, Joshua, et al.
Published: (2025)
by: Ward, Joshua, et al.
Published: (2025)
When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs
by: Shen, Xinyue, et al.
Published: (2025)
by: Shen, Xinyue, et al.
Published: (2025)
Similar Items
-
Asking Forever: Universal Activations Behind Turn Amplification in Conversational LLMs
by: Coalson, Zachary, et al.
Published: (2026) -
Fail-Closed Alignment for Large Language Models
by: Coalson, Zachary, et al.
Published: (2026) -
Hard Work Does Not Always Pay Off: Poisoning Attacks on Neural Architecture Search
by: Coalson, Zachary, et al.
Published: (2024) -
IF-GUIDE: Influence Function-Guided Detoxification of LLMs
by: Coalson, Zachary, et al.
Published: (2025) -
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
by: Wen, Yuxin, et al.
Published: (2024)