Adaptive PII Mitigation Framework for Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Asthana, Shubhi, Mahindru, Ruchi, Zhang, Bing, Sanz, Jorge |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Deploying Privacy Guardrails for LLMs: A Comparative Analysis of Real-World Applications
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)
Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models
di: Sivashanmugam, Sathesh P.
Pubblicazione: (2025)
di: Sivashanmugam, Sathesh P.
Pubblicazione: (2025)
STRIDE: A Systematic Framework for Selecting AI Modalities -- Agentic AI, AI Assistants, or LLM Calls
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
di: Peng, Benji, et al.
Pubblicazione: (2024)
di: Peng, Benji, et al.
Pubblicazione: (2024)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
di: Nakka, Krishna Kanth, et al.
Pubblicazione: (2024)
di: Nakka, Krishna Kanth, et al.
Pubblicazione: (2024)
Oblivionis: A Lightweight Learning and Unlearning Framework for Federated Large Language Models
di: Zhang, Fuyao, et al.
Pubblicazione: (2025)
di: Zhang, Fuyao, et al.
Pubblicazione: (2025)
LISAA: A Framework for Large Language Model Information Security Awareness Assessment
di: Cohen, Ofir, et al.
Pubblicazione: (2024)
di: Cohen, Ofir, et al.
Pubblicazione: (2024)
Large Language Model Empowered Privacy-Protected Framework for PHI Annotation in Clinical Notes
di: Wu, Guanchen, et al.
Pubblicazione: (2025)
di: Wu, Guanchen, et al.
Pubblicazione: (2025)
LMEraser: Large Model Unlearning through Adaptive Prompt Tuning
di: Xu, Jie, et al.
Pubblicazione: (2024)
di: Xu, Jie, et al.
Pubblicazione: (2024)
Information Theoretic Adversarial Training of Large Language Models
di: Zhang, Yiwei, et al.
Pubblicazione: (2026)
di: Zhang, Yiwei, et al.
Pubblicazione: (2026)
Privacy Auditing of Large Language Models
di: Panda, Ashwinee, et al.
Pubblicazione: (2025)
di: Panda, Ashwinee, et al.
Pubblicazione: (2025)
Watermark Stealing in Large Language Models
di: Jovanović, Nikola, et al.
Pubblicazione: (2024)
di: Jovanović, Nikola, et al.
Pubblicazione: (2024)
DMark: Order-Agnostic Watermarking for Diffusion Large Language Models
di: Wu, Linyu, et al.
Pubblicazione: (2025)
di: Wu, Linyu, et al.
Pubblicazione: (2025)
Model-based Large Language Model Customization as Service
di: Wu, Zhaomin, et al.
Pubblicazione: (2024)
di: Wu, Zhaomin, et al.
Pubblicazione: (2024)
Exploring the Secondary Risks of Large Language Models
di: Chen, Jiawei, et al.
Pubblicazione: (2025)
di: Chen, Jiawei, et al.
Pubblicazione: (2025)
Finetuning Large Language Models for Vulnerability Detection
di: Shestov, Alexey, et al.
Pubblicazione: (2024)
di: Shestov, Alexey, et al.
Pubblicazione: (2024)
Research on Key Technologies for Cross-Cloud Federated Training of Large Language Models
di: Yang, Haowei, et al.
Pubblicazione: (2024)
di: Yang, Haowei, et al.
Pubblicazione: (2024)
Generalist++: A Meta-learning Framework for Mitigating Trade-off in Adversarial Training
di: Wang, Yisen, et al.
Pubblicazione: (2025)
di: Wang, Yisen, et al.
Pubblicazione: (2025)
UpSafe$^\circ$C: Upcycling for Controllable Safety in Large Language Models
di: Sun, Yuhao, et al.
Pubblicazione: (2025)
di: Sun, Yuhao, et al.
Pubblicazione: (2025)
Large Language Models Are Unreliable for Cyber Threat Intelligence
di: Mezzi, Emanuele, et al.
Pubblicazione: (2025)
di: Mezzi, Emanuele, et al.
Pubblicazione: (2025)
Prompt Injection Attacks on Large Language Models in Oncology
di: Clusmann, Jan, et al.
Pubblicazione: (2024)
di: Clusmann, Jan, et al.
Pubblicazione: (2024)
Towards Characterizing Cyber Networks with Large Language Models
di: Hartsock, Alaric, et al.
Pubblicazione: (2024)
di: Hartsock, Alaric, et al.
Pubblicazione: (2024)
Understanding the Effects of Safety Unalignment on Large Language Models
di: Halloran, John T.
Pubblicazione: (2026)
di: Halloran, John T.
Pubblicazione: (2026)
Mitigating Watermark Forgery in Generative Models via Randomized Key Selection
di: Aremu, Toluwani, et al.
Pubblicazione: (2025)
di: Aremu, Toluwani, et al.
Pubblicazione: (2025)
A Survey on Model Extraction Attacks and Defenses for Large Language Models
di: Zhao, Kaixiang, et al.
Pubblicazione: (2025)
di: Zhao, Kaixiang, et al.
Pubblicazione: (2025)
CTIGuardian: A Few-Shot Framework for Mitigating Privacy Leakage in Fine-Tuned LLMs
di: Arachchige, Shashie Dilhara Batan, et al.
Pubblicazione: (2025)
di: Arachchige, Shashie Dilhara Batan, et al.
Pubblicazione: (2025)
Permissioned LLMs: Enforcing Access Control in Large Language Models
di: Jayaraman, Bargav, et al.
Pubblicazione: (2025)
di: Jayaraman, Bargav, et al.
Pubblicazione: (2025)
Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques
di: Jaffal, Niveen O., et al.
Pubblicazione: (2025)
di: Jaffal, Niveen O., et al.
Pubblicazione: (2025)
Securing Large Language Models: Threats, Vulnerabilities and Responsible Practices
di: Abdali, Sara, et al.
Pubblicazione: (2024)
di: Abdali, Sara, et al.
Pubblicazione: (2024)
Evaluating Large Language Models for Security Bug Report Prediction
di: Soltaniani, Farnaz, et al.
Pubblicazione: (2026)
di: Soltaniani, Farnaz, et al.
Pubblicazione: (2026)
WebPII: Benchmarking Visual PII Detection for Computer-Use Agents
di: Zhao, Nathan
Pubblicazione: (2026)
di: Zhao, Nathan
Pubblicazione: (2026)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
di: Liu, Guozhi, et al.
Pubblicazione: (2025)
di: Liu, Guozhi, et al.
Pubblicazione: (2025)
Large Language Models in Wireless Application Design: In-Context Learning-enhanced Automatic Network Intrusion Detection
di: Zhang, Han, et al.
Pubblicazione: (2024)
di: Zhang, Han, et al.
Pubblicazione: (2024)
DR-Encoder: Encode Low-rank Gradients with Random Prior for Large Language Models Differentially Privately
di: Wu, Huiwen, et al.
Pubblicazione: (2024)
di: Wu, Huiwen, et al.
Pubblicazione: (2024)
Beyond Data Privacy: New Privacy Risks for Large Language Models
di: Du, Yuntao, et al.
Pubblicazione: (2025)
di: Du, Yuntao, et al.
Pubblicazione: (2025)
Evaluating the Reliability and Fidelity of Automated Judgment Systems of Large Language Models
di: Biskupski, Tom, et al.
Pubblicazione: (2026)
di: Biskupski, Tom, et al.
Pubblicazione: (2026)
Differentially Private Preference Data Synthesis for Large Language Model Alignment
di: Gao, Fengyu, et al.
Pubblicazione: (2026)
di: Gao, Fengyu, et al.
Pubblicazione: (2026)
Reconstruction of Differentially Private Text Sanitization via Large Language Models
di: Pang, Shuchao, et al.
Pubblicazione: (2024)
di: Pang, Shuchao, et al.
Pubblicazione: (2024)
Improved Large Language Model Jailbreak Detection via Pretrained Embeddings
di: Galinkin, Erick, et al.
Pubblicazione: (2024)
di: Galinkin, Erick, et al.
Pubblicazione: (2024)
PLeak: Prompt Leaking Attacks against Large Language Model Applications
di: Hui, Bo, et al.
Pubblicazione: (2024)
di: Hui, Bo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Deploying Privacy Guardrails for LLMs: A Comparative Analysis of Real-World Applications
di: Asthana, Shubhi, et al.
Pubblicazione: (2025) -
Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models
di: Sivashanmugam, Sathesh P.
Pubblicazione: (2025) -
STRIDE: A Systematic Framework for Selecting AI Modalities -- Agentic AI, AI Assistants, or LLM Calls
di: Asthana, Shubhi, et al.
Pubblicazione: (2025) -
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
di: Peng, Benji, et al.
Pubblicazione: (2024) -
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
di: Nakka, Krishna Kanth, et al.
Pubblicazione: (2024)