Privately Aligning Language Models with Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Fan, Inan, Huseyin A., Backurs, Arturs, Chandrasekaran, Varun, Kulkarni, Janardhan, Sim, Robert |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Differentially Private Training of Mixture of Experts Models
by: Tholoniat, Pierre, et al.
Published: (2024)
by: Tholoniat, Pierre, et al.
Published: (2024)
Privacy-Preserving In-Context Learning with Differentially Private Few-Shot Generation
by: Tang, Xinyu, et al.
Published: (2023)
by: Tang, Xinyu, et al.
Published: (2023)
Efficiently Computing Similarities to Private Datasets
by: Backurs, Arturs, et al.
Published: (2024)
by: Backurs, Arturs, et al.
Published: (2024)
Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models
by: Wang, Xiangwen, et al.
Published: (2026)
by: Wang, Xiangwen, et al.
Published: (2026)
Challenges in Enabling Private Data Valuation
by: Fu, Yiwei, et al.
Published: (2026)
by: Fu, Yiwei, et al.
Published: (2026)
Selective Pre-training for Private Fine-tuning
by: Yu, Da, et al.
Published: (2023)
by: Yu, Da, et al.
Published: (2023)
Differentially Private Synthetic Data via Foundation Model APIs 1: Images
by: Lin, Zinan, et al.
Published: (2023)
by: Lin, Zinan, et al.
Published: (2023)
AMUN: Adversarial Machine UNlearning
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2025)
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2025)
Bypassing LLM Watermarks with Color-Aware Substitutions
by: Wu, Qilong, et al.
Published: (2024)
by: Wu, Qilong, et al.
Published: (2024)
Individual Privacy Accounting for Differentially Private Stochastic Gradient Descent
by: Yu, Da, et al.
Published: (2022)
by: Yu, Da, et al.
Published: (2022)
Practical Differentially Private Hyperparameter Tuning with Subsampling
by: Koskela, Antti, et al.
Published: (2023)
by: Koskela, Antti, et al.
Published: (2023)
LanFL: Differentially Private Federated Learning with Large Language Models using Synthetic Samples
by: Wu, Huiyu, et al.
Published: (2024)
by: Wu, Huiyu, et al.
Published: (2024)
Adversarial Attacks on Locally Private Graph Neural Networks
by: Varun, Matta, et al.
Published: (2026)
by: Varun, Matta, et al.
Published: (2026)
Differentially Private In-Context Learning with Nearest Neighbor Search
by: Koskela, Antti, et al.
Published: (2025)
by: Koskela, Antti, et al.
Published: (2025)
The Erasure Illusion: Stress-Testing the Generalization of LLM Forgetting Evaluation
by: Jia, Hengrui, et al.
Published: (2025)
by: Jia, Hengrui, et al.
Published: (2025)
Scaling Laws for Differentially Private Language Models
by: McKenna, Ryan, et al.
Published: (2025)
by: McKenna, Ryan, et al.
Published: (2025)
PrivORL: Differentially Private Synthetic Dataset for Offline Reinforcement Learning
by: Gong, Chen, et al.
Published: (2025)
by: Gong, Chen, et al.
Published: (2025)
Differentially Private Deep Model-Based Reinforcement Learning
by: Rio, Alexandre, et al.
Published: (2024)
by: Rio, Alexandre, et al.
Published: (2024)
Temporal Context Awareness: A Defense Framework Against Multi-turn Manipulation Attacks on Large Language Models
by: Kulkarni, Prashant, et al.
Published: (2025)
by: Kulkarni, Prashant, et al.
Published: (2025)
Differentially Private Low-Rank Adaptation of Large Language Model Using Federated Learning
by: Liu, Xiao-Yang, et al.
Published: (2023)
by: Liu, Xiao-Yang, et al.
Published: (2023)
DP-SelFT: Differentially Private Selective Fine-Tuning for Large Language Models
by: Sha, Haichao, et al.
Published: (2026)
by: Sha, Haichao, et al.
Published: (2026)
Distributional Black-Box Model Inversion Attack with Multi-Agent Reinforcement Learning
by: Bao, Huan, et al.
Published: (2024)
by: Bao, Huan, et al.
Published: (2024)
Bayesian Pseudo Posterior Mechanism for Differentially Private Machine Learning
by: Chew, Robert, et al.
Published: (2025)
by: Chew, Robert, et al.
Published: (2025)
Differentially Private Subspace Fine-Tuning for Large Language Models
by: Zheng, Lele, et al.
Published: (2026)
by: Zheng, Lele, et al.
Published: (2026)
Adaptively Private Next-Token Prediction of Large Language Models
by: Flemings, James, et al.
Published: (2024)
by: Flemings, James, et al.
Published: (2024)
Oracle-Efficient Differentially Private Learning with Public Data
by: Block, Adam, et al.
Published: (2024)
by: Block, Adam, et al.
Published: (2024)
Privately Learning Decision Lists and a Differentially Private Winnow
by: Bun, Mark, et al.
Published: (2026)
by: Bun, Mark, et al.
Published: (2026)
Private and Communication-Efficient Federated Learning based on Differentially Private Sketches
by: Zhang, Meifan, et al.
Published: (2024)
by: Zhang, Meifan, et al.
Published: (2024)
XAI and Android Malware Models
by: Kulkarni, Maithili, et al.
Published: (2024)
by: Kulkarni, Maithili, et al.
Published: (2024)
The Efficacy of Transfer-based No-box Attacks on Image Watermarking: A Pragmatic Analysis
by: Wu, Qilong, et al.
Published: (2024)
by: Wu, Qilong, et al.
Published: (2024)
FlashDP: Private Training Large Language Models with Efficient DP-SGD
by: Wang, Liangyu, et al.
Published: (2025)
by: Wang, Liangyu, et al.
Published: (2025)
On the Benefits of Public Representations for Private Transfer Learning under Distribution Shift
by: Thaker, Pratiksha, et al.
Published: (2023)
by: Thaker, Pratiksha, et al.
Published: (2023)
Learning with Locally Private Examples by Inverse Weierstrass Private Stochastic Gradient Descent
by: Dufraiche, Jean, et al.
Published: (2026)
by: Dufraiche, Jean, et al.
Published: (2026)
Differentially Private Diffusion Models
by: Dockhorn, Tim, et al.
Published: (2022)
by: Dockhorn, Tim, et al.
Published: (2022)
TMI! Finetuned Models Leak Private Information from their Pretraining Data
by: Abascal, John, et al.
Published: (2023)
by: Abascal, John, et al.
Published: (2023)
Correlated Noise Mechanisms for Differentially Private Learning
by: Pillutla, Krishna, et al.
Published: (2025)
by: Pillutla, Krishna, et al.
Published: (2025)
Layer-Targeted Multilingual Knowledge Erasure in Large Language Models
by: Li, Taoran, et al.
Published: (2026)
by: Li, Taoran, et al.
Published: (2026)
Mitigating Noise Detriment in Differentially Private Federated Learning with Model Pre-training
by: Jin, Huitong, et al.
Published: (2024)
by: Jin, Huitong, et al.
Published: (2024)
Differentially Private Random Feature Model
by: Liao, Chunyang, et al.
Published: (2024)
by: Liao, Chunyang, et al.
Published: (2024)
Naturally Private Recommendations with Determinantal Point Processes
by: Fitzsimons, Jack, et al.
Published: (2024)
by: Fitzsimons, Jack, et al.
Published: (2024)
Similar Items
-
Differentially Private Training of Mixture of Experts Models
by: Tholoniat, Pierre, et al.
Published: (2024) -
Privacy-Preserving In-Context Learning with Differentially Private Few-Shot Generation
by: Tang, Xinyu, et al.
Published: (2023) -
Efficiently Computing Similarities to Private Datasets
by: Backurs, Arturs, et al.
Published: (2024) -
Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models
by: Wang, Xiangwen, et al.
Published: (2026) -
Challenges in Enabling Private Data Valuation
by: Fu, Yiwei, et al.
Published: (2026)