Failure Modes of Maximum Entropy RLHF
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Çağatan, Ömer Veysel, Akgün, Barış |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uncovering RL Integration in SSL Loss: Objective-Specific Implications for Data-Efficient RL
von: Çağatan, Ömer Veysel, et al.
Veröffentlicht: (2024)
von: Çağatan, Ömer Veysel, et al.
Veröffentlicht: (2024)
SigCLR: Sigmoid Contrastive Learning of Visual Representations
von: Çağatan, Ömer Veysel
Veröffentlicht: (2024)
von: Çağatan, Ömer Veysel
Veröffentlicht: (2024)
Higher Resolution, Better Generalization: Unlocking Visual Scaling in Deep Reinforcement Learning
von: Trumpp, Raphael, et al.
Veröffentlicht: (2026)
von: Trumpp, Raphael, et al.
Veröffentlicht: (2026)
UNSEE: Unsupervised Non-contrastive Sentence Embeddings
von: Çağatan, Ömer Veysel
Veröffentlicht: (2024)
von: Çağatan, Ömer Veysel
Veröffentlicht: (2024)
Clipping-Free Policy Optimization for Large Language Models
von: Çağatan, Ömer Veysel, et al.
Veröffentlicht: (2026)
von: Çağatan, Ömer Veysel, et al.
Veröffentlicht: (2026)
Failure Modes of LLMs for Causal Reasoning on Narratives
von: Yamin, Khurram, et al.
Veröffentlicht: (2024)
von: Yamin, Khurram, et al.
Veröffentlicht: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
von: Hu, Jian, et al.
Veröffentlicht: (2024)
von: Hu, Jian, et al.
Veröffentlicht: (2024)
Evaluating Defences against Unsafe Feedback in RLHF
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
Solving the Inverse Alignment Problem for Efficient RLHF
von: Krishna, Shambhavi, et al.
Veröffentlicht: (2024)
von: Krishna, Shambhavi, et al.
Veröffentlicht: (2024)
Group Robust Preference Optimization in Reward-free RLHF
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2024)
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2024)
Why Is RLHF Alignment Shallow? A Gradient Analysis
von: Young, Robin
Veröffentlicht: (2026)
von: Young, Robin
Veröffentlicht: (2026)
RLHF and IIA: Perverse Incentives
von: Xu, Wanqiao, et al.
Veröffentlicht: (2023)
von: Xu, Wanqiao, et al.
Veröffentlicht: (2023)
Reward-Robust RLHF in LLMs
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
A Long Way to Go: Investigating Length Correlations in RLHF
von: Singhal, Prasann, et al.
Veröffentlicht: (2023)
von: Singhal, Prasann, et al.
Veröffentlicht: (2023)
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
von: Shi, Ruizhe, et al.
Veröffentlicht: (2025)
von: Shi, Ruizhe, et al.
Veröffentlicht: (2025)
Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
Reward Model Overoptimisation in Iterated RLHF
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024)
von: Frick, Evan, et al.
Veröffentlicht: (2024)
Dataset Reset Policy Optimization for RLHF
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
Measuring memorization in RLHF for code completion
von: Pappu, Aneesh, et al.
Veröffentlicht: (2024)
von: Pappu, Aneesh, et al.
Veröffentlicht: (2024)
Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
von: Fu, Yuqian, et al.
Veröffentlicht: (2026)
von: Fu, Yuqian, et al.
Veröffentlicht: (2026)
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
von: Pal, Arka, et al.
Veröffentlicht: (2024)
von: Pal, Arka, et al.
Veröffentlicht: (2024)
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
Information-Theoretic Reward Decomposition for Generalizable RLHF
von: Mao, Liyuan, et al.
Veröffentlicht: (2025)
von: Mao, Liyuan, et al.
Veröffentlicht: (2025)
General Exploratory Bonus for Optimistic Exploration in RLHF
von: Li, Wendi, et al.
Veröffentlicht: (2025)
von: Li, Wendi, et al.
Veröffentlicht: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
Quantile Regression for Distributional Reward Models in RLHF
von: Dorka, Nicolai
Veröffentlicht: (2024)
von: Dorka, Nicolai
Veröffentlicht: (2024)
The Perfect Blend: Redefining RLHF with Mixture of Judges
von: Xu, Tengyu, et al.
Veröffentlicht: (2024)
von: Xu, Tengyu, et al.
Veröffentlicht: (2024)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Understanding the Effects of RLHF on LLM Generalisation and Diversity
von: Kirk, Robert, et al.
Veröffentlicht: (2023)
von: Kirk, Robert, et al.
Veröffentlicht: (2023)
WPO: Enhancing RLHF with Weighted Preference Optimization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
Advancing Translation Preference Modeling with RLHF: A Step Towards Cost-Effective Solution
von: Xu, Nuo, et al.
Veröffentlicht: (2024)
von: Xu, Nuo, et al.
Veröffentlicht: (2024)
Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization
von: Yao, Jiashu, et al.
Veröffentlicht: (2026)
von: Yao, Jiashu, et al.
Veröffentlicht: (2026)
Adversarial Robustness of Discriminative Self-Supervised Learning in Vision
von: Çağatan, Ömer Veysel, et al.
Veröffentlicht: (2025)
von: Çağatan, Ömer Veysel, et al.
Veröffentlicht: (2025)
Adaptive Margin RLHF via Preference over Preferences
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
DPO Meets PPO: Reinforced Token Optimization for RLHF
von: Zhong, Han, et al.
Veröffentlicht: (2024)
von: Zhong, Han, et al.
Veröffentlicht: (2024)
An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training
von: Xiao, Youshao, et al.
Veröffentlicht: (2023)
von: Xiao, Youshao, et al.
Veröffentlicht: (2023)
Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models
von: Dam, Harvey, et al.
Veröffentlicht: (2025)
von: Dam, Harvey, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Uncovering RL Integration in SSL Loss: Objective-Specific Implications for Data-Efficient RL
von: Çağatan, Ömer Veysel, et al.
Veröffentlicht: (2024) -
SigCLR: Sigmoid Contrastive Learning of Visual Representations
von: Çağatan, Ömer Veysel
Veröffentlicht: (2024) -
Higher Resolution, Better Generalization: Unlocking Visual Scaling in Deep Reinforcement Learning
von: Trumpp, Raphael, et al.
Veröffentlicht: (2026) -
UNSEE: Unsupervised Non-contrastive Sentence Embeddings
von: Çağatan, Ömer Veysel
Veröffentlicht: (2024) -
Clipping-Free Policy Optimization for Large Language Models
von: Çağatan, Ömer Veysel, et al.
Veröffentlicht: (2026)