RAPO: Risk-Aware Preference Optimization for Generalizable Safe Reasoning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wei, Zeming, Zhang, Qiaosheng, Hu, Xia, Xu, Xingcheng |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Boosting Jailbreak Attack with Momentum
par: Zhang, Yihao, et autres
Publié: (2024)
par: Zhang, Yihao, et autres
Publié: (2024)
Secure LLM Fine-Tuning via Safety-Aware Probing
par: Wu, Chengcan, et autres
Publié: (2025)
par: Wu, Chengcan, et autres
Publié: (2025)
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
par: Zhang, Zhixin, et autres
Publié: (2025)
par: Zhang, Zhixin, et autres
Publié: (2025)
Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
par: Zhang, Yihao, et autres
Publié: (2024)
par: Zhang, Yihao, et autres
Publié: (2024)
On the Duality Between Sharpness-Aware Minimization and Adversarial Training
par: Zhang, Yihao, et autres
Publié: (2024)
par: Zhang, Yihao, et autres
Publié: (2024)
Exploring the Robustness of In-Context Learning with Noisy Labels
par: Cheng, Chen, et autres
Publié: (2024)
par: Cheng, Chen, et autres
Publié: (2024)
Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation
par: Wu, Chengcan, et autres
Publié: (2025)
par: Wu, Chengcan, et autres
Publié: (2025)
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
par: Jin, Chang, et autres
Publié: (2026)
par: Jin, Chang, et autres
Publié: (2026)
Identifying and Understanding Cross-Class Features in Adversarial Training
par: Wei, Zeming, et autres
Publié: (2025)
par: Wei, Zeming, et autres
Publié: (2025)
Calibrated Adversarial Sampling: Multi-Armed Bandit-Guided Generalization Against Unforeseen Attacks
par: Wang, Rui, et autres
Publié: (2025)
par: Wang, Rui, et autres
Publié: (2025)
A Queueing-Theoretic Framework for Dynamic Attack Surfaces: Data-Integrated Risk Analysis and Adaptive Defense
par: Yun, Jihyeon, et autres
Publié: (2026)
par: Yun, Jihyeon, et autres
Publié: (2026)
Clip-and-Verify: Linear Constraint-Driven Domain Clipping for Accelerating Neural Network Verification
par: Zhou, Duo, et autres
Publié: (2025)
par: Zhou, Duo, et autres
Publié: (2025)
Approximate and Weighted Data Reconstruction Attack in Federated Learning
par: Song, Yongcun, et autres
Publié: (2023)
par: Song, Yongcun, et autres
Publié: (2023)
Characterizing the Training Dynamics of Private Fine-tuning with Langevin diffusion
par: Ke, Shuqi, et autres
Publié: (2024)
par: Ke, Shuqi, et autres
Publié: (2024)
Correlated Noise Provably Beats Independent Noise for Differentially Private Learning
par: Choquette-Choo, Christopher A., et autres
Publié: (2023)
par: Choquette-Choo, Christopher A., et autres
Publié: (2023)
SMI: Statistical Membership Inference for Reliable Unlearned Model Auditing
par: Sun, Jialong, et autres
Publié: (2026)
par: Sun, Jialong, et autres
Publié: (2026)
On Model Protection in Federated Learning against Eavesdropping Attacks
par: Maity, Dipankar, et autres
Publié: (2025)
par: Maity, Dipankar, et autres
Publié: (2025)
TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems
par: Wang, Kai, et autres
Publié: (2026)
par: Wang, Kai, et autres
Publié: (2026)
Private Zeroth-Order Nonsmooth Nonconvex Optimization
par: Zhang, Qinzi, et autres
Publié: (2024)
par: Zhang, Qinzi, et autres
Publié: (2024)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
par: Mo, Yichuan, et autres
Publié: (2024)
par: Mo, Yichuan, et autres
Publié: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
par: Chen, Taiye, et autres
Publié: (2025)
par: Chen, Taiye, et autres
Publié: (2025)
The Power of Sampling: Dimension-free Risk Bounds in Private ERM
par: Lee, Yin Tat, et autres
Publié: (2021)
par: Lee, Yin Tat, et autres
Publié: (2021)
Differentially Private Bilevel Optimization
par: Kornowski, Guy
Publié: (2024)
par: Kornowski, Guy
Publié: (2024)
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
par: Wei, Zeming, et autres
Publié: (2026)
par: Wei, Zeming, et autres
Publié: (2026)
Differential Privacy via Distributionally Robust Optimization
par: Selvi, Aras, et autres
Publié: (2023)
par: Selvi, Aras, et autres
Publié: (2023)
On the Gradient Complexity of Private Optimization with Private Oracles
par: Menart, Michael, et autres
Publié: (2025)
par: Menart, Michael, et autres
Publié: (2025)
Efficient Optimization Algorithms for Linear Adversarial Training
par: RIbeiro, Antônio H., et autres
Publié: (2024)
par: RIbeiro, Antônio H., et autres
Publié: (2024)
Improved Sample Complexity for Private Nonsmooth Nonconvex Optimization
par: Kornowski, Guy, et autres
Publié: (2024)
par: Kornowski, Guy, et autres
Publié: (2024)
Public-data Assisted Private Stochastic Optimization: Power and Limitations
par: Ullah, Enayat, et autres
Publié: (2024)
par: Ullah, Enayat, et autres
Publié: (2024)
Private Stochastic Optimization With Large Worst-Case Lipschitz Parameter
par: Lowy, Andrew, et autres
Publié: (2022)
par: Lowy, Andrew, et autres
Publié: (2022)
Faster Algorithms for User-Level Private Stochastic Convex Optimization
par: Lowy, Andrew, et autres
Publié: (2024)
par: Lowy, Andrew, et autres
Publié: (2024)
Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions
par: Islamov, Rustem, et autres
Publié: (2026)
par: Islamov, Rustem, et autres
Publié: (2026)
Differentially Private Bilevel Optimization: Efficient Algorithms with Near-Optimal Rates
par: Lowy, Andrew, et autres
Publié: (2025)
par: Lowy, Andrew, et autres
Publié: (2025)
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
par: Wei, Zeming, et autres
Publié: (2023)
par: Wei, Zeming, et autres
Publié: (2023)
Differentially Private Non-Convex Optimization under the KL Condition with Optimal Rates
par: Menart, Michael, et autres
Publié: (2023)
par: Menart, Michael, et autres
Publié: (2023)
MILE: A Mutation Testing Framework of In-Context Learning Systems
par: Wei, Zeming, et autres
Publié: (2024)
par: Wei, Zeming, et autres
Publié: (2024)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
par: He, Luxi, et autres
Publié: (2024)
par: He, Luxi, et autres
Publié: (2024)
How to Make the Gradients Small Privately: Improved Rates for Differentially Private Non-Convex Optimization
par: Lowy, Andrew, et autres
Publié: (2024)
par: Lowy, Andrew, et autres
Publié: (2024)
Automata-Based Steering of Large Language Models for Diverse Structured Generation
par: Luan, Xiaokun, et autres
Publié: (2025)
par: Luan, Xiaokun, et autres
Publié: (2025)
Scalable Neural Network Verification with Branch-and-bound Inferred Cutting Planes
par: Zhou, Duo, et autres
Publié: (2024)
par: Zhou, Duo, et autres
Publié: (2024)
Documents similaires
-
Boosting Jailbreak Attack with Momentum
par: Zhang, Yihao, et autres
Publié: (2024) -
Secure LLM Fine-Tuning via Safety-Aware Probing
par: Wu, Chengcan, et autres
Publié: (2025) -
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
par: Zhang, Zhixin, et autres
Publié: (2025) -
Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
par: Zhang, Yihao, et autres
Publié: (2024) -
On the Duality Between Sharpness-Aware Minimization and Adversarial Training
par: Zhang, Yihao, et autres
Publié: (2024)