On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
Fuente:
arXiv
Saved in:
| Main Authors: | Nishimori, Soichiro, Zhang, Yu-Jie, Lodkaew, Thanawat, Sugiyama, Masashi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
by: Ono, Shinnosuke, et al.
Published: (2026)
by: Ono, Shinnosuke, et al.
Published: (2026)
Mahjax: A GPU-Accelerated Mahjong Simulator for Reinforcement Learning in JAX
by: Nishimori, Soichiro, et al.
Published: (2026)
by: Nishimori, Soichiro, et al.
Published: (2026)
RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences
by: Cheng, Jie, et al.
Published: (2024)
by: Cheng, Jie, et al.
Published: (2024)
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
by: Nishimori, Soichiro, et al.
Published: (2026)
by: Nishimori, Soichiro, et al.
Published: (2026)
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
by: Cai, Xin-Qiang, et al.
Published: (2025)
by: Cai, Xin-Qiang, et al.
Published: (2025)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
by: Ackermann, Johannes, et al.
Published: (2025)
by: Ackermann, Johannes, et al.
Published: (2025)
Offline Reinforcement Learning with Domain-Unlabeled Data
by: Nishimori, Soichiro, et al.
Published: (2024)
by: Nishimori, Soichiro, et al.
Published: (2024)
Sharpness-Aware Black-Box Optimization
by: Ye, Feiyang, et al.
Published: (2024)
by: Ye, Feiyang, et al.
Published: (2024)
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction
by: Cai, Xin-Qiang, et al.
Published: (2026)
by: Cai, Xin-Qiang, et al.
Published: (2026)
Pgx: Hardware-Accelerated Parallel Game Simulators for Reinforcement Learning
by: Koyamada, Sotetsu, et al.
Published: (2023)
by: Koyamada, Sotetsu, et al.
Published: (2023)
Recursive Reward Aggregation
by: Tang, Yuting, et al.
Published: (2025)
by: Tang, Yuting, et al.
Published: (2025)
Symmetric Behavior Regularized Policy Optimization
by: Zhu, Lingwei, et al.
Published: (2025)
by: Zhu, Lingwei, et al.
Published: (2025)
Learning Robust Diffusion Models from Imprecise Supervision
by: Wu, Dong-Dong, et al.
Published: (2025)
by: Wu, Dong-Dong, et al.
Published: (2025)
Offline Reinforcement Learning from Datasets with Structured Non-Stationarity
by: Ackermann, Johannes, et al.
Published: (2024)
by: Ackermann, Johannes, et al.
Published: (2024)
Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated Perturbation
by: Xu, Jie, et al.
Published: (2025)
by: Xu, Jie, et al.
Published: (2025)
Symmetric Reinforcement Learning Loss for Robust Learning on Diverse Tasks and Model Scales
by: Byun, Ju-Seung, et al.
Published: (2024)
by: Byun, Ju-Seung, et al.
Published: (2024)
Breaking the Curse of Repulsion: Optimistic Distributionally Robust Policy Optimization for Off-Policy Generative Recommendation
by: Jiang, Jie, et al.
Published: (2026)
by: Jiang, Jie, et al.
Published: (2026)
Self-Improving Robust Preference Optimization
by: Choi, Eugene, et al.
Published: (2024)
by: Choi, Eugene, et al.
Published: (2024)
Mitigating Preference Hacking in Policy Optimization with Pessimism
by: Gupta, Dhawal, et al.
Published: (2025)
by: Gupta, Dhawal, et al.
Published: (2025)
On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm Perspective
by: Xie, Zeke, et al.
Published: (2020)
by: Xie, Zeke, et al.
Published: (2020)
Towards Scalable Oversight via Partitioned Human Supervision
by: Yin, Ren, et al.
Published: (2025)
by: Yin, Ren, et al.
Published: (2025)
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
by: Zhang, Tianle, et al.
Published: (2024)
by: Zhang, Tianle, et al.
Published: (2024)
Bootstrapping LLMs via Preference-Based Policy Optimization
by: Jia, Chen
Published: (2025)
by: Jia, Chen
Published: (2025)
From Noisy Traces to Stable Gradients: Bias-Variance Optimized Preference Optimization for Aligning Large Reasoning Models
by: Zhu, Mingkang, et al.
Published: (2025)
by: Zhu, Mingkang, et al.
Published: (2025)
Preference Learning with Response Time: Robust Losses and Guarantees
by: Sawarni, Ayush, et al.
Published: (2025)
by: Sawarni, Ayush, et al.
Published: (2025)
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
by: Zhou, Zhanhui, et al.
Published: (2023)
by: Zhou, Zhanhui, et al.
Published: (2023)
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
by: Xu, Zaiyan, et al.
Published: (2025)
by: Xu, Zaiyan, et al.
Published: (2025)
Adversarial Policy Optimization for Offline Preference-based Reinforcement Learning
by: Kang, Hyungkyu, et al.
Published: (2025)
by: Kang, Hyungkyu, et al.
Published: (2025)
Optimal Transport for LLM Reward Modeling from Noisy Preference
by: Pan, Licheng, et al.
Published: (2026)
by: Pan, Licheng, et al.
Published: (2026)
ROPO: Robust Preference Optimization for Large Language Models
by: Liang, Xize, et al.
Published: (2024)
by: Liang, Xize, et al.
Published: (2024)
Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought
by: Zhang, Zhen-Yu, et al.
Published: (2024)
by: Zhang, Zhen-Yu, et al.
Published: (2024)
ANO: A Principled Approach to Robust Policy Optimization
by: Zhang, Yiheng, et al.
Published: (2026)
by: Zhang, Yiheng, et al.
Published: (2026)
Impact of Noisy Supervision in Foundation Model Learning
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
From Coefficients to Directions: Rethinking Model Merging with Directional Alignment
by: Chen, Zhikang, et al.
Published: (2025)
by: Chen, Zhikang, et al.
Published: (2025)
TGPO: Tree-Guided Preference Optimization for Robust Web Agent Reinforcement Learning
by: Chen, Ziyuan, et al.
Published: (2025)
by: Chen, Ziyuan, et al.
Published: (2025)
Matrix Sensing with Kernel Optimal Loss: Robustness and Optimization Landscape
by: Song, Xinyuan, et al.
Published: (2025)
by: Song, Xinyuan, et al.
Published: (2025)
Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation
by: Nakanishi, Kosuke, et al.
Published: (2025)
by: Nakanishi, Kosuke, et al.
Published: (2025)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
by: Ackermann, Johannes, et al.
Published: (2026)
by: Ackermann, Johannes, et al.
Published: (2026)
COPR: Continual Human Preference Learning via Optimal Policy Regularization
by: Zhang, Han, et al.
Published: (2024)
by: Zhang, Han, et al.
Published: (2024)
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
by: Cho, Taehyun, et al.
Published: (2025)
by: Cho, Taehyun, et al.
Published: (2025)
Similar Items
-
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
by: Ono, Shinnosuke, et al.
Published: (2026) -
Mahjax: A GPU-Accelerated Mahjong Simulator for Reinforcement Learning in JAX
by: Nishimori, Soichiro, et al.
Published: (2026) -
RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences
by: Cheng, Jie, et al.
Published: (2024) -
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
by: Nishimori, Soichiro, et al.
Published: (2026) -
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
by: Cai, Xin-Qiang, et al.
Published: (2025)