RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
Fuente:
arXiv
Salvato in:
| Autori principali: | Liang, Kaiqu, Hu, Haimin, Liu, Ryan, Griffiths, Thomas L., Fisac, Jaime Fernández |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
di: Liang, Kaiqu, et al.
Pubblicazione: (2025)
di: Liang, Kaiqu, et al.
Pubblicazione: (2025)
Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
di: Liang, Kaiqu, et al.
Pubblicazione: (2024)
di: Liang, Kaiqu, et al.
Pubblicazione: (2024)
Learning Personalized Agents from Human Feedback
di: Liang, Kaiqu, et al.
Pubblicazione: (2026)
di: Liang, Kaiqu, et al.
Pubblicazione: (2026)
Reward Shaping to Mitigate Reward Hacking in RLHF
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
di: Chen, Lichang, et al.
Pubblicazione: (2024)
di: Chen, Lichang, et al.
Pubblicazione: (2024)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
di: Zhu, Banghua, et al.
Pubblicazione: (2024)
di: Zhu, Banghua, et al.
Pubblicazione: (2024)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
di: Ono, Shinnosuke, et al.
Pubblicazione: (2026)
di: Ono, Shinnosuke, et al.
Pubblicazione: (2026)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
di: Hu, Jian, et al.
Pubblicazione: (2024)
di: Hu, Jian, et al.
Pubblicazione: (2024)
Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
di: Hahm, Dongyoon, et al.
Pubblicazione: (2025)
di: Hahm, Dongyoon, et al.
Pubblicazione: (2025)
Provable Interactive Learning with Hindsight Instruction Feedback
di: Misra, Dipendra, et al.
Pubblicazione: (2024)
di: Misra, Dipendra, et al.
Pubblicazione: (2024)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
di: Hu, Michael Y., et al.
Pubblicazione: (2025)
di: Hu, Michael Y., et al.
Pubblicazione: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
di: Dong, Hanze, et al.
Pubblicazione: (2024)
di: Dong, Hanze, et al.
Pubblicazione: (2024)
How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?
di: Liu, Ryan, et al.
Pubblicazione: (2024)
di: Liu, Ryan, et al.
Pubblicazione: (2024)
CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
di: Butt, Natasha, et al.
Pubblicazione: (2024)
di: Butt, Natasha, et al.
Pubblicazione: (2024)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
di: Hong, Joey, et al.
Pubblicazione: (2024)
di: Hong, Joey, et al.
Pubblicazione: (2024)
An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training
di: Xiao, Youshao, et al.
Pubblicazione: (2023)
di: Xiao, Youshao, et al.
Pubblicazione: (2023)
HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
di: Yeo, Woongyeng, et al.
Pubblicazione: (2026)
di: Yeo, Woongyeng, et al.
Pubblicazione: (2026)
RLHF and IIA: Perverse Incentives
di: Xu, Wanqiao, et al.
Pubblicazione: (2023)
di: Xu, Wanqiao, et al.
Pubblicazione: (2023)
Reward-Robust RLHF in LLMs
di: Yan, Yuzi, et al.
Pubblicazione: (2024)
di: Yan, Yuzi, et al.
Pubblicazione: (2024)
MAGICS: Adversarial RL with Minimax Actors Guided by Implicit Critic Stackelberg for Convergent Neural Synthesis of Robot Safety
di: Wang, Justin, et al.
Pubblicazione: (2024)
di: Wang, Justin, et al.
Pubblicazione: (2024)
Are Large Language Models Sensitive to the Motives Behind Communication?
di: Wu, Addison J., et al.
Pubblicazione: (2025)
di: Wu, Addison J., et al.
Pubblicazione: (2025)
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
di: Wu, Yuning, et al.
Pubblicazione: (2026)
di: Wu, Yuning, et al.
Pubblicazione: (2026)
HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning
di: Lu, Zhicong, et al.
Pubblicazione: (2026)
di: Lu, Zhicong, et al.
Pubblicazione: (2026)
Reward Model Overoptimisation in Iterated RLHF
di: Wolf, Lorenz, et al.
Pubblicazione: (2025)
di: Wolf, Lorenz, et al.
Pubblicazione: (2025)
How to Evaluate Reward Models for RLHF
di: Frick, Evan, et al.
Pubblicazione: (2024)
di: Frick, Evan, et al.
Pubblicazione: (2024)
Dataset Reset Policy Optimization for RLHF
di: Chang, Jonathan D., et al.
Pubblicazione: (2024)
di: Chang, Jonathan D., et al.
Pubblicazione: (2024)
Epistemic Traps: Rational Misalignment Driven by Model Misspecification
di: Xu, Xingcheng, et al.
Pubblicazione: (2026)
di: Xu, Xingcheng, et al.
Pubblicazione: (2026)
Information-Theoretic Reward Decomposition for Generalizable RLHF
di: Mao, Liyuan, et al.
Pubblicazione: (2025)
di: Mao, Liyuan, et al.
Pubblicazione: (2025)
General Exploratory Bonus for Optimistic Exploration in RLHF
di: Li, Wendi, et al.
Pubblicazione: (2025)
di: Li, Wendi, et al.
Pubblicazione: (2025)
Active Preference Optimization for Sample Efficient RLHF
di: Das, Nirjhar, et al.
Pubblicazione: (2024)
di: Das, Nirjhar, et al.
Pubblicazione: (2024)
Quantile Regression for Distributional Reward Models in RLHF
di: Dorka, Nicolai
Pubblicazione: (2024)
di: Dorka, Nicolai
Pubblicazione: (2024)
The Perfect Blend: Redefining RLHF with Mixture of Judges
di: Xu, Tengyu, et al.
Pubblicazione: (2024)
di: Xu, Tengyu, et al.
Pubblicazione: (2024)
Understanding the Effects of RLHF on LLM Generalisation and Diversity
di: Kirk, Robert, et al.
Pubblicazione: (2023)
di: Kirk, Robert, et al.
Pubblicazione: (2023)
WPO: Enhancing RLHF with Weighted Preference Optimization
di: Zhou, Wenxuan, et al.
Pubblicazione: (2024)
di: Zhou, Wenxuan, et al.
Pubblicazione: (2024)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
di: Zhu, Yu, et al.
Pubblicazione: (2024)
di: Zhu, Yu, et al.
Pubblicazione: (2024)
Large Language Models Assume People are More Rational than We Really are
di: Liu, Ryan, et al.
Pubblicazione: (2024)
di: Liu, Ryan, et al.
Pubblicazione: (2024)
Identifying and Mitigating the Influence of the Prior Distribution in Large Language Models
di: Zhang, Liyi, et al.
Pubblicazione: (2025)
di: Zhang, Liyi, et al.
Pubblicazione: (2025)
Adaptive Margin RLHF via Preference over Preferences
di: Chittepu, Yaswanth, et al.
Pubblicazione: (2025)
di: Chittepu, Yaswanth, et al.
Pubblicazione: (2025)
DPO Meets PPO: Reinforced Token Optimization for RLHF
di: Zhong, Han, et al.
Pubblicazione: (2024)
di: Zhong, Han, et al.
Pubblicazione: (2024)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
di: Lee, Harrison, et al.
Pubblicazione: (2023)
di: Lee, Harrison, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
di: Liang, Kaiqu, et al.
Pubblicazione: (2025) -
Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
di: Liang, Kaiqu, et al.
Pubblicazione: (2024) -
Learning Personalized Agents from Human Feedback
di: Liang, Kaiqu, et al.
Pubblicazione: (2026) -
Reward Shaping to Mitigate Reward Hacking in RLHF
di: Fu, Jiayi, et al.
Pubblicazione: (2025) -
ODIN: Disentangled Reward Mitigates Hacking in RLHF
di: Chen, Lichang, et al.
Pubblicazione: (2024)