Pr$εε$mpt: Sanitizing Sensitive Prompts for LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Chowdhury, Amrita Roy, Glukhov, David, Anshumaan, Divyam, Chalasani, Prasad, Papernot, Nicolas, Jha, Somesh, Bellare, Mihir |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Not to Detect Prompt Injections with an LLM
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
Dependency-Aware Privacy for Multi-turn Agents
di: Anshumaan, Divyam, et al.
Pubblicazione: (2026)
di: Anshumaan, Divyam, et al.
Pubblicazione: (2026)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
di: Wang, Zi, et al.
Pubblicazione: (2024)
di: Wang, Zi, et al.
Pubblicazione: (2024)
PolicyLR: A Logic Representation For Privacy Policies
di: Hooda, Ashish, et al.
Pubblicazione: (2024)
di: Hooda, Ashish, et al.
Pubblicazione: (2024)
Augment then Smooth: Reconciling Differential Privacy with Certified Robustness
di: Wu, Jiapeng, et al.
Pubblicazione: (2023)
di: Wu, Jiapeng, et al.
Pubblicazione: (2023)
What Really is a Member? Discrediting Membership Inference via Poisoning
di: Mangaokar, Neal, et al.
Pubblicazione: (2025)
di: Mangaokar, Neal, et al.
Pubblicazione: (2025)
Formal Policy Enforcement for Real-World Agentic Systems
di: Palumbo, Nils, et al.
Pubblicazione: (2026)
di: Palumbo, Nils, et al.
Pubblicazione: (2026)
Breach By A Thousand Leaks: Unsafe Information Leakage in `Safe' AI Responses
di: Glukhov, David, et al.
Pubblicazione: (2024)
di: Glukhov, David, et al.
Pubblicazione: (2024)
Beyond Laplace and Gaussian: Exploring the Generalized Gaussian Mechanism for Private Machine Learning
di: Rinberg, Roy, et al.
Pubblicazione: (2025)
di: Rinberg, Roy, et al.
Pubblicazione: (2025)
Tighter Privacy Auditing of DP-SGD in the Hidden State Threat Model
di: Cebere, Tudor, et al.
Pubblicazione: (2024)
di: Cebere, Tudor, et al.
Pubblicazione: (2024)
Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
di: Thudi, Anvith, et al.
Pubblicazione: (2023)
di: Thudi, Anvith, et al.
Pubblicazione: (2023)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
di: Geng, Runpeng, et al.
Pubblicazione: (2025)
di: Geng, Runpeng, et al.
Pubblicazione: (2025)
Fast Exact Unlearning for In-Context Learning Data for LLMs
di: Muresanu, Andrei I., et al.
Pubblicazione: (2024)
di: Muresanu, Andrei I., et al.
Pubblicazione: (2024)
Have it your way: Individualized Privacy Assignment for DP-SGD
di: Boenisch, Franziska, et al.
Pubblicazione: (2023)
di: Boenisch, Franziska, et al.
Pubblicazione: (2023)
Beyond Labeling Oracles: What does it mean to steal ML models?
di: Shafran, Avital, et al.
Pubblicazione: (2023)
di: Shafran, Avital, et al.
Pubblicazione: (2023)
SLVR: Securely Leveraging Client Validation for Robust Federated Learning
di: Choi, Jihye, et al.
Pubblicazione: (2025)
di: Choi, Jihye, et al.
Pubblicazione: (2025)
On the Difficulty of Constructing a Robust and Publicly-Detectable Watermark
di: Fairoze, Jaiden, et al.
Pubblicazione: (2025)
di: Fairoze, Jaiden, et al.
Pubblicazione: (2025)
FairProof : Confidential and Certifiable Fairness for Neural Networks
di: Yadav, Chhavi, et al.
Pubblicazione: (2024)
di: Yadav, Chhavi, et al.
Pubblicazione: (2024)
On the Privacy Risk of In-context Learning
di: Duan, Haonan, et al.
Pubblicazione: (2024)
di: Duan, Haonan, et al.
Pubblicazione: (2024)
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
di: Hayes, Jamie, et al.
Pubblicazione: (2024)
di: Hayes, Jamie, et al.
Pubblicazione: (2024)
Private Rate-Constrained Optimization with Applications to Fair Learning
di: Yaghini, Mohammad, et al.
Pubblicazione: (2025)
di: Yaghini, Mohammad, et al.
Pubblicazione: (2025)
When Vision Fails: Text Attacks Against ViT and OCR
di: Boucher, Nicholas, et al.
Pubblicazione: (2023)
di: Boucher, Nicholas, et al.
Pubblicazione: (2023)
Robustness of Locally Differentially Private Graph Analysis Against Poisoning
di: Imola, Jacob, et al.
Pubblicazione: (2022)
di: Imola, Jacob, et al.
Pubblicazione: (2022)
Metric Differential Privacy at the User-Level Via the Earth Mover's Distance
di: Imola, Jacob, et al.
Pubblicazione: (2024)
di: Imola, Jacob, et al.
Pubblicazione: (2024)
Confused ChatGPT: Cross-App Context Poisoning via First-Party APIs
di: Wang, Chao, et al.
Pubblicazione: (2026)
di: Wang, Chao, et al.
Pubblicazione: (2026)
CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
di: Das, Debeshee, et al.
Pubblicazione: (2025)
di: Das, Debeshee, et al.
Pubblicazione: (2025)
The Fundamental Limits of Least-Privilege Learning
di: Stadler, Theresa, et al.
Pubblicazione: (2024)
di: Stadler, Theresa, et al.
Pubblicazione: (2024)
Learned-Database Systems Security
di: Schuster, Roei, et al.
Pubblicazione: (2022)
di: Schuster, Roei, et al.
Pubblicazione: (2022)
LLM Dataset Inference: Did you train on my dataset?
di: Maini, Pratyush, et al.
Pubblicazione: (2024)
di: Maini, Pratyush, et al.
Pubblicazione: (2024)
PURL: Safe and Effective Sanitization of Link Decoration
di: Munir, Shaoor, et al.
Pubblicazione: (2023)
di: Munir, Shaoor, et al.
Pubblicazione: (2023)
LLMStinger: Jailbreaking LLMs using RL fine-tuned LLMs
di: Jha, Piyush, et al.
Pubblicazione: (2024)
di: Jha, Piyush, et al.
Pubblicazione: (2024)
Efficient Public Verification of Private ML via Regularization
di: Bell, Zoë Ruha, et al.
Pubblicazione: (2025)
di: Bell, Zoë Ruha, et al.
Pubblicazione: (2025)
Entropy-Guided Attention for Private LLMs
di: Jha, Nandan Kumar, et al.
Pubblicazione: (2025)
di: Jha, Nandan Kumar, et al.
Pubblicazione: (2025)
Publicly-Detectable Watermarking for Language Models
di: Fairoze, Jaiden, et al.
Pubblicazione: (2023)
di: Fairoze, Jaiden, et al.
Pubblicazione: (2023)
Decentralised, Collaborative, and Privacy-preserving Machine Learning for Multi-Hospital Data
di: Fang, Congyu, et al.
Pubblicazione: (2024)
di: Fang, Congyu, et al.
Pubblicazione: (2024)
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
di: Liu, Xiaogeng, et al.
Pubblicazione: (2024)
di: Liu, Xiaogeng, et al.
Pubblicazione: (2024)
Preempting Text Sanitization Utility in Resource-Constrained Privacy-Preserving LLM Interactions
di: Carpentier, Robin, et al.
Pubblicazione: (2024)
di: Carpentier, Robin, et al.
Pubblicazione: (2024)
FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization
di: Zhu, Shunan, et al.
Pubblicazione: (2026)
di: Zhu, Shunan, et al.
Pubblicazione: (2026)
Backdoor Detection through Replicated Execution of Outsourced Training
di: Jia, Hengrui, et al.
Pubblicazione: (2025)
di: Jia, Hengrui, et al.
Pubblicazione: (2025)
Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs
di: Agarwal, Krishiv, et al.
Pubblicazione: (2026)
di: Agarwal, Krishiv, et al.
Pubblicazione: (2026)
Documenti analoghi
-
How Not to Detect Prompt Injections with an LLM
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025) -
Dependency-Aware Privacy for Multi-turn Agents
di: Anshumaan, Divyam, et al.
Pubblicazione: (2026) -
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
di: Wang, Zi, et al.
Pubblicazione: (2024) -
PolicyLR: A Logic Representation For Privacy Policies
di: Hooda, Ashish, et al.
Pubblicazione: (2024) -
Augment then Smooth: Reconciling Differential Privacy with Certified Robustness
di: Wu, Jiapeng, et al.
Pubblicazione: (2023)