Improved Bounds for Private and Robust Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Weng, Wenqian, He, Yi, Zhou, Xingyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Sample Complexity of Differentially Private Policy Optimization
by: He, Yi, et al.
Published: (2025)
by: He, Yi, et al.
Published: (2025)
Towards Differentially Private Reinforcement Learning with General Function Approximation
by: He, Yi, et al.
Published: (2026)
by: He, Yi, et al.
Published: (2026)
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
by: Zhou, Xingyu, et al.
Published: (2025)
by: Zhou, Xingyu, et al.
Published: (2025)
Square$χ$PO: Differentially Private and Robust $χ^2$-Preference Optimization in Offline Direct Alignment
by: Zhou, Xingyu, et al.
Published: (2025)
by: Zhou, Xingyu, et al.
Published: (2025)
Private Wasserstein Distance
by: Li, Wenqian, et al.
Published: (2024)
by: Li, Wenqian, et al.
Published: (2024)
Improved Algorithms for Differentially Private Language Model Alignment
by: Chen, Keyu, et al.
Published: (2025)
by: Chen, Keyu, et al.
Published: (2025)
Doubly Robust Alignment for Large Language Models
by: Xu, Erhan, et al.
Published: (2025)
by: Xu, Erhan, et al.
Published: (2025)
Intelligent Icing Detection Model of Wind Turbine Blades Based on SCADA data
by: Jiang, Wenqian, et al.
Published: (2021)
by: Jiang, Wenqian, et al.
Published: (2021)
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
by: Shen, Sicheng, et al.
Published: (2026)
by: Shen, Sicheng, et al.
Published: (2026)
Exploring the Impact of Temperature Scaling in Softmax for Classification and Adversarial Robustness
by: Xuan, Hao, et al.
Published: (2025)
by: Xuan, Hao, et al.
Published: (2025)
Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs
by: Ao, Shuang, et al.
Published: (2025)
by: Ao, Shuang, et al.
Published: (2025)
Provably Robust Conformal Prediction with Improved Efficiency
by: Yan, Ge, et al.
Published: (2024)
by: Yan, Ge, et al.
Published: (2024)
Conformal Feedback Alignment: Quantifying Answer-Level Reliability for Robust LLM Alignment
by: Chen, Tiejin, et al.
Published: (2026)
by: Chen, Tiejin, et al.
Published: (2026)
Symbolic Graph Networks for Robust PDE Discovery from Noisy Sparse Data
by: Chen, Xingyu, et al.
Published: (2026)
by: Chen, Xingyu, et al.
Published: (2026)
Tabular Data Adapters: Improving Outlier Detection for Unlabeled Private Data
by: Herurkar, Dayananda, et al.
Published: (2025)
by: Herurkar, Dayananda, et al.
Published: (2025)
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
by: Xu, Zaiyan, et al.
Published: (2025)
by: Xu, Zaiyan, et al.
Published: (2025)
Adversarial Preference Learning for Robust LLM Alignment
by: Wang, Yuanfu, et al.
Published: (2025)
by: Wang, Yuanfu, et al.
Published: (2025)
Manifold Approximation leads to Robust Kernel Alignment
by: Islam, Mohammad Tariqul, et al.
Published: (2025)
by: Islam, Mohammad Tariqul, et al.
Published: (2025)
Weight Ensembling Improves Reasoning in Language Models
by: Dang, Xingyu, et al.
Published: (2025)
by: Dang, Xingyu, et al.
Published: (2025)
Rectifying Shortcut Behaviors in Preference-based Reward Learning
by: Ye, Wenqian, et al.
Published: (2025)
by: Ye, Wenqian, et al.
Published: (2025)
Dynamic Search for Inference-Time Alignment in Diffusion Models
by: Li, Xiner, et al.
Published: (2025)
by: Li, Xiner, et al.
Published: (2025)
Auditing Approximate Machine Unlearning for Differentially Private Models
by: Gu, Yuechun, et al.
Published: (2025)
by: Gu, Yuechun, et al.
Published: (2025)
Tight Lower Bounds and Improved Convergence in Performative Prediction
by: Khorsandi, Pedram, et al.
Published: (2024)
by: Khorsandi, Pedram, et al.
Published: (2024)
Improved Generalization Bounds for Communication Efficient Federated Learning
by: Gholami, Peyman, et al.
Published: (2024)
by: Gholami, Peyman, et al.
Published: (2024)
The Cost of Robustness: Tighter Bounds on Parameter Complexity for Robust Memorization in ReLU Nets
by: Kim, Yujun, et al.
Published: (2025)
by: Kim, Yujun, et al.
Published: (2025)
Learning Adaptive Distribution Alignment with Neural Characteristic Function for Graph Domain Adaptation
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
Breaking the Barrier: Enhanced Utility and Robustness in Smoothed DRL Agents
by: Sun, Chung-En, et al.
Published: (2024)
by: Sun, Chung-En, et al.
Published: (2024)
Rethinking Inverse Reinforcement Learning: from Data Alignment to Task Alignment
by: Zhou, Weichao, et al.
Published: (2024)
by: Zhou, Weichao, et al.
Published: (2024)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
by: Guo, Hanze, et al.
Published: (2025)
by: Guo, Hanze, et al.
Published: (2025)
Differentially Private Preference Data Synthesis for Large Language Model Alignment
by: Gao, Fengyu, et al.
Published: (2026)
by: Gao, Fengyu, et al.
Published: (2026)
Efficient Robust Conformal Prediction via Lipschitz-Bounded Networks
by: Massena, Thomas, et al.
Published: (2025)
by: Massena, Thomas, et al.
Published: (2025)
Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning
by: Moradipari, Ahmadreza, et al.
Published: (2023)
by: Moradipari, Ahmadreza, et al.
Published: (2023)
Self-adaptive weights based on balanced residual decay rate for physics-informed neural networks and deep operator networks
by: Chen, Wenqian, et al.
Published: (2024)
by: Chen, Wenqian, et al.
Published: (2024)
Noisy-Pair Robust Representation Alignment for Positive-Unlabeled Learning
by: Zhao, Hengwei, et al.
Published: (2025)
by: Zhao, Hengwei, et al.
Published: (2025)
Self-Improving Robust Preference Optimization
by: Choi, Eugene, et al.
Published: (2024)
by: Choi, Eugene, et al.
Published: (2024)
Upper and Lower Bounds for Distributionally Robust Off-Dynamics Reinforcement Learning
by: Liu, Zhishuai, et al.
Published: (2024)
by: Liu, Zhishuai, et al.
Published: (2024)
Linearizing Models for Efficient yet Robust Private Inference
by: Sarkar, Sreetama, et al.
Published: (2024)
by: Sarkar, Sreetama, et al.
Published: (2024)
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
by: Rajaram, Sara, et al.
Published: (2025)
by: Rajaram, Sara, et al.
Published: (2025)
Scalable Valuation of Human Feedback through Provably Robust Model Alignment
by: Fujisawa, Masahiro, et al.
Published: (2025)
by: Fujisawa, Masahiro, et al.
Published: (2025)
Similar Items
-
On the Sample Complexity of Differentially Private Policy Optimization
by: He, Yi, et al.
Published: (2025) -
Towards Differentially Private Reinforcement Learning with General Function Approximation
by: He, Yi, et al.
Published: (2026) -
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
by: Zhou, Xingyu, et al.
Published: (2025) -
Square$χ$PO: Differentially Private and Robust $χ^2$-Preference Optimization in Offline Direct Alignment
by: Zhou, Xingyu, et al.
Published: (2025) -
Private Wasserstein Distance
by: Li, Wenqian, et al.
Published: (2024)