CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Mireshghallah, Niloofar, Mangaokar, Neal, Kokhlikyan, Narine, Zharmagambetov, Arman, Zaheer, Manzil, Mahloujifar, Saeed, Chaudhuri, Kamalika |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
di: Wen, Yuxin, et al.
Pubblicazione: (2025)
di: Wen, Yuxin, et al.
Pubblicazione: (2025)
SecAlign: Defending Against Prompt Injection with Preference Optimization
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
Auditing $f$-Differential Privacy in One Run
di: Mahloujifar, Saeed, et al.
Pubblicazione: (2024)
di: Mahloujifar, Saeed, et al.
Pubblicazione: (2024)
Z0-Inf: Zeroth Order Approximation for Data Influence
di: Kokhlikyan, Narine, et al.
Pubblicazione: (2025)
di: Kokhlikyan, Narine, et al.
Pubblicazione: (2025)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)
Machine Learning with Privacy for Protected Attributes
di: Mahloujifar, Saeed, et al.
Pubblicazione: (2025)
di: Mahloujifar, Saeed, et al.
Pubblicazione: (2025)
Privacy Amplification for the Gaussian Mechanism via Bounded Support
di: Hu, Shengyuan, et al.
Pubblicazione: (2024)
di: Hu, Shengyuan, et al.
Pubblicazione: (2024)
Privacy Blur: Quantifying Privacy and Utility for Image Data Release
di: Mahloujifar, Saeed, et al.
Pubblicazione: (2025)
di: Mahloujifar, Saeed, et al.
Pubblicazione: (2025)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023)
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023)
Guarantees of confidentiality via Hammersley-Chapman-Robbins bounds
di: Chaudhuri, Kamalika, et al.
Pubblicazione: (2024)
di: Chaudhuri, Kamalika, et al.
Pubblicazione: (2024)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
di: Naseh, Ali, et al.
Pubblicazione: (2025)
di: Naseh, Ali, et al.
Pubblicazione: (2025)
Measuring Privacy Loss in Distributed Spatio-Temporal Data
di: Koga, Tatsuki, et al.
Pubblicazione: (2024)
di: Koga, Tatsuki, et al.
Pubblicazione: (2024)
Differentially Private Model Merging
di: Yin, Qichuan, et al.
Pubblicazione: (2026)
di: Yin, Qichuan, et al.
Pubblicazione: (2026)
Position: Privacy Is Not Just Memorization!
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2025)
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2025)
Detecting Benchmark Contamination Through Watermarking
di: Sander, Tom, et al.
Pubblicazione: (2025)
di: Sander, Tom, et al.
Pubblicazione: (2025)
Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
Robustness of Locally Differentially Private Graph Analysis Against Poisoning
di: Imola, Jacob, et al.
Pubblicazione: (2022)
di: Imola, Jacob, et al.
Pubblicazione: (2022)
Metric Differential Privacy at the User-Level Via the Earth Mover's Distance
di: Imola, Jacob, et al.
Pubblicazione: (2024)
di: Imola, Jacob, et al.
Pubblicazione: (2024)
Communication-Efficient Triangle Counting under Local Differential Privacy
di: Imola, Jacob, et al.
Pubblicazione: (2021)
di: Imola, Jacob, et al.
Pubblicazione: (2021)
Can We Infer Confidential Properties of Training Data from LLMs?
di: Huang, Pengrun, et al.
Pubblicazione: (2025)
di: Huang, Pengrun, et al.
Pubblicazione: (2025)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
di: Wang, Erchi, et al.
Pubblicazione: (2026)
di: Wang, Erchi, et al.
Pubblicazione: (2026)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
di: Paulus, Anselm, et al.
Pubblicazione: (2024)
di: Paulus, Anselm, et al.
Pubblicazione: (2024)
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
di: Fu, Wenjie, et al.
Pubblicazione: (2026)
di: Fu, Wenjie, et al.
Pubblicazione: (2026)
Observational Auditing of Label Privacy
di: Kalemaj, Iden, et al.
Pubblicazione: (2025)
di: Kalemaj, Iden, et al.
Pubblicazione: (2025)
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
di: Mangaokar, Neal, et al.
Pubblicazione: (2024)
di: Mangaokar, Neal, et al.
Pubblicazione: (2024)
Integrating Differential Privacy and Contextual Integrity
di: Benthall, Sebastian, et al.
Pubblicazione: (2024)
di: Benthall, Sebastian, et al.
Pubblicazione: (2024)
It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
di: Park, Sangwoo, et al.
Pubblicazione: (2026)
di: Park, Sangwoo, et al.
Pubblicazione: (2026)
Defending Object Detectors against Patch Attacks with Out-of-Distribution Smoothing
di: Feng, Ryan, et al.
Pubblicazione: (2022)
di: Feng, Ryan, et al.
Pubblicazione: (2022)
Differentially Private Learning Needs Better Model Initialization and Self-Distillation
di: Ngong, Ivoline C., et al.
Pubblicazione: (2024)
di: Ngong, Ivoline C., et al.
Pubblicazione: (2024)
Can Large Language Models Really Recognize Your Name?
di: Pham, Dzung, et al.
Pubblicazione: (2025)
di: Pham, Dzung, et al.
Pubblicazione: (2025)
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
di: Panda, Ashwinee, et al.
Pubblicazione: (2022)
di: Panda, Ashwinee, et al.
Pubblicazione: (2022)
Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy
di: Koga, Tatsuki, et al.
Pubblicazione: (2024)
di: Koga, Tatsuki, et al.
Pubblicazione: (2024)
PPMI: Privacy-Preserving LLM Interaction with Socratic Chain-of-Thought Reasoning and Homomorphically Encrypted Vector Databases
di: Bae, Yubeen, et al.
Pubblicazione: (2025)
di: Bae, Yubeen, et al.
Pubblicazione: (2025)
Private Fine-tuning of Large Language Models with Zeroth-order Optimization
di: Tang, Xinyu, et al.
Pubblicazione: (2024)
di: Tang, Xinyu, et al.
Pubblicazione: (2024)
Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training
di: Borkar, Jaydeep, et al.
Pubblicazione: (2025)
di: Borkar, Jaydeep, et al.
Pubblicazione: (2025)
What Really is a Member? Discrediting Membership Inference via Poisoning
di: Mangaokar, Neal, et al.
Pubblicazione: (2025)
di: Mangaokar, Neal, et al.
Pubblicazione: (2025)
FairProof : Confidential and Certifiable Fairness for Neural Networks
di: Yadav, Chhavi, et al.
Pubblicazione: (2024)
di: Yadav, Chhavi, et al.
Pubblicazione: (2024)
On Differentially Private U Statistics
di: Chaudhuri, Kamalika, et al.
Pubblicazione: (2024)
di: Chaudhuri, Kamalika, et al.
Pubblicazione: (2024)
ExpProof : Operationalizing Explanations for Confidential Models with ZKPs
di: Yadav, Chhavi, et al.
Pubblicazione: (2025)
di: Yadav, Chhavi, et al.
Pubblicazione: (2025)
Taming Data Challenges in ML-based Security Tasks Using Generative AI
di: Kanchi, Shravya, et al.
Pubblicazione: (2025)
di: Kanchi, Shravya, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
di: Wen, Yuxin, et al.
Pubblicazione: (2025) -
SecAlign: Defending Against Prompt Injection with Preference Optimization
di: Chen, Sizhe, et al.
Pubblicazione: (2024) -
Auditing $f$-Differential Privacy in One Run
di: Mahloujifar, Saeed, et al.
Pubblicazione: (2024) -
Z0-Inf: Zeroth Order Approximation for Data Influence
di: Kokhlikyan, Narine, et al.
Pubblicazione: (2025) -
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)