Can We Infer Confidential Properties of Training Data from LLMs?
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Pengrun, Yadav, Chhavi, Chaudhuri, Kamalika, Wu, Ruihan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ExpProof : Operationalizing Explanations for Confidential Models with ZKPs
di: Yadav, Chhavi, et al.
Pubblicazione: (2025)
di: Yadav, Chhavi, et al.
Pubblicazione: (2025)
FairProof : Confidential and Certifiable Fairness for Neural Networks
di: Yadav, Chhavi, et al.
Pubblicazione: (2024)
di: Yadav, Chhavi, et al.
Pubblicazione: (2024)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
di: Wang, Erchi, et al.
Pubblicazione: (2026)
di: Wang, Erchi, et al.
Pubblicazione: (2026)
Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy
di: Koga, Tatsuki, et al.
Pubblicazione: (2024)
di: Koga, Tatsuki, et al.
Pubblicazione: (2024)
Better Membership Inference Privacy Measurement through Discrepancy
di: Wu, Ruihan, et al.
Pubblicazione: (2024)
di: Wu, Ruihan, et al.
Pubblicazione: (2024)
Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
di: Guo, Wenkai, et al.
Pubblicazione: (2025)
di: Guo, Wenkai, et al.
Pubblicazione: (2025)
Auditing $f$-Differential Privacy in One Run
di: Mahloujifar, Saeed, et al.
Pubblicazione: (2024)
di: Mahloujifar, Saeed, et al.
Pubblicazione: (2024)
Learning-Time Encoding Shapes Unlearning in LLMs
di: Wu, Ruihan, et al.
Pubblicazione: (2025)
di: Wu, Ruihan, et al.
Pubblicazione: (2025)
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
di: Wallace, Eric, et al.
Pubblicazione: (2024)
di: Wallace, Eric, et al.
Pubblicazione: (2024)
Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks
di: Kuo, Kevin, et al.
Pubblicazione: (2026)
di: Kuo, Kevin, et al.
Pubblicazione: (2026)
Influence-based Attributions can be Manipulated
di: Yadav, Chhavi, et al.
Pubblicazione: (2024)
di: Yadav, Chhavi, et al.
Pubblicazione: (2024)
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
di: Yoon, Do-hyeon, et al.
Pubblicazione: (2025)
di: Yoon, Do-hyeon, et al.
Pubblicazione: (2025)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
di: Naseh, Ali, et al.
Pubblicazione: (2025)
di: Naseh, Ali, et al.
Pubblicazione: (2025)
Learning from Negative Examples: Why Warning-Framed Training Data Teaches What It Warns Against
di: Enkhbayar, Tsogt-Ochir
Pubblicazione: (2025)
di: Enkhbayar, Tsogt-Ochir
Pubblicazione: (2025)
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
di: Wen, Yuxin, et al.
Pubblicazione: (2025)
di: Wen, Yuxin, et al.
Pubblicazione: (2025)
Machine Learning with Privacy for Protected Attributes
di: Mahloujifar, Saeed, et al.
Pubblicazione: (2025)
di: Mahloujifar, Saeed, et al.
Pubblicazione: (2025)
Tracing Privacy Leakage of Language Models to Training Data via Adjusted Influence Functions
di: Liu, Jinxin, et al.
Pubblicazione: (2024)
di: Liu, Jinxin, et al.
Pubblicazione: (2024)
Guarantees of confidentiality via Hammersley-Chapman-Robbins bounds
di: Chaudhuri, Kamalika, et al.
Pubblicazione: (2024)
di: Chaudhuri, Kamalika, et al.
Pubblicazione: (2024)
Inferring Properties of Graph Neural Networks
di: Nguyen, Dat, et al.
Pubblicazione: (2024)
di: Nguyen, Dat, et al.
Pubblicazione: (2024)
A Curious Case of Searching for the Correlation between Training Data and Adversarial Robustness of Transformer Textual Models
di: Dang, Cuong, et al.
Pubblicazione: (2024)
di: Dang, Cuong, et al.
Pubblicazione: (2024)
Privacy Amplification for the Gaussian Mechanism via Bounded Support
di: Hu, Shengyuan, et al.
Pubblicazione: (2024)
di: Hu, Shengyuan, et al.
Pubblicazione: (2024)
Detecting Pretraining Data from Large Language Models
di: Shi, Weijia, et al.
Pubblicazione: (2023)
di: Shi, Weijia, et al.
Pubblicazione: (2023)
Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content
di: Liu, Bing, et al.
Pubblicazione: (2026)
di: Liu, Bing, et al.
Pubblicazione: (2026)
How Vulnerable Are Edge LLMs?
di: Ding, Ao, et al.
Pubblicazione: (2026)
di: Ding, Ao, et al.
Pubblicazione: (2026)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
di: Suri, Omar Farooq Khan, et al.
Pubblicazione: (2025)
di: Suri, Omar Farooq Khan, et al.
Pubblicazione: (2025)
Adaptive Pre-training Data Detection for Large Language Models via Surprising Tokens
di: Zhang, Anqi, et al.
Pubblicazione: (2024)
di: Zhang, Anqi, et al.
Pubblicazione: (2024)
Promoting Data and Model Privacy in Federated Learning through Quantized LoRA
di: Zhu, JianHao, et al.
Pubblicazione: (2024)
di: Zhu, JianHao, et al.
Pubblicazione: (2024)
HARMONIC: Harnessing LLMs for Tabular Data Synthesis and Privacy Protection
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
ContinuousBench: Can Differentially Private Synthetic Text Improve Capabilities?
di: Liu, Peihan, et al.
Pubblicazione: (2026)
di: Liu, Peihan, et al.
Pubblicazione: (2026)
UCD: Unlearning in LLMs via Contrastive Decoding
di: Suriyakumar, Vinith M., et al.
Pubblicazione: (2025)
di: Suriyakumar, Vinith M., et al.
Pubblicazione: (2025)
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
di: Dotsinski, Asen, et al.
Pubblicazione: (2026)
di: Dotsinski, Asen, et al.
Pubblicazione: (2026)
Coercing LLMs to do and reveal (almost) anything
di: Geiping, Jonas, et al.
Pubblicazione: (2024)
di: Geiping, Jonas, et al.
Pubblicazione: (2024)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
di: Vega, Jason, et al.
Pubblicazione: (2023)
di: Vega, Jason, et al.
Pubblicazione: (2023)
LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTI
di: Schwartz, Yuval, et al.
Pubblicazione: (2024)
di: Schwartz, Yuval, et al.
Pubblicazione: (2024)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
di: Sivapiromrat, Sanhanat, et al.
Pubblicazione: (2025)
di: Sivapiromrat, Sanhanat, et al.
Pubblicazione: (2025)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
di: Zhao, Xuandong, et al.
Pubblicazione: (2024)
di: Zhao, Xuandong, et al.
Pubblicazione: (2024)
Evaluating Deep Unlearning in Large Language Models
di: Wu, Ruihan, et al.
Pubblicazione: (2024)
di: Wu, Ruihan, et al.
Pubblicazione: (2024)
Prompt Public Large Language Models to Synthesize Data for Private On-device Applications
di: Wu, Shanshan, et al.
Pubblicazione: (2024)
di: Wu, Shanshan, et al.
Pubblicazione: (2024)
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
di: Wang, Linlin, et al.
Pubblicazione: (2025)
di: Wang, Linlin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ExpProof : Operationalizing Explanations for Confidential Models with ZKPs
di: Yadav, Chhavi, et al.
Pubblicazione: (2025) -
FairProof : Confidential and Certifiable Fairness for Neural Networks
di: Yadav, Chhavi, et al.
Pubblicazione: (2024) -
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
di: Wang, Erchi, et al.
Pubblicazione: (2026) -
Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy
di: Koga, Tatsuki, et al.
Pubblicazione: (2024) -
Better Membership Inference Privacy Measurement through Discrepancy
di: Wu, Ruihan, et al.
Pubblicazione: (2024)