SAFE: Stable Alignment Finetuning with Entropy-Aware Predictive Control for Reinforcement Learning from Human Feedback (RLHF)
Fuente:
arXiv
Salvato in:
| Autore principale: | Maity, Dipan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AuON: A Linear-time Alternative to Orthogonal Momentum Updates
di: Maity, Dipan
Pubblicazione: (2025)
di: Maity, Dipan
Pubblicazione: (2025)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
di: Lee, Harrison, et al.
Pubblicazione: (2023)
di: Lee, Harrison, et al.
Pubblicazione: (2023)
Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
di: Ji, Jiaming, et al.
Pubblicazione: (2025)
di: Ji, Jiaming, et al.
Pubblicazione: (2025)
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
di: Chaudhari, Shreyas, et al.
Pubblicazione: (2024)
di: Chaudhari, Shreyas, et al.
Pubblicazione: (2024)
Gated-SwinRMT: Unifying Swin Windowed Attention with Retentive Manhattan Decay via Input-Dependent Gating
di: Maity, Dipan, et al.
Pubblicazione: (2026)
di: Maity, Dipan, et al.
Pubblicazione: (2026)
ACE-RLHF: Automated Code Evaluation and Socratic Feedback Generation Tool using Large Language Models and Reinforcement Learning with Human Feedback
di: Rahman, Tasnia, et al.
Pubblicazione: (2025)
di: Rahman, Tasnia, et al.
Pubblicazione: (2025)
The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback
di: Lambert, Nathan, et al.
Pubblicazione: (2023)
di: Lambert, Nathan, et al.
Pubblicazione: (2023)
Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback
di: Yuan, Yifu, et al.
Pubblicazione: (2024)
di: Yuan, Yifu, et al.
Pubblicazione: (2024)
RLHF Fine-Tuning of LLMs for Alignment with Implicit User Feedback in Conversational Recommenders
di: Yang, Zhongheng, et al.
Pubblicazione: (2025)
di: Yang, Zhongheng, et al.
Pubblicazione: (2025)
SAFE-RL: Saliency-Aware Counterfactual Explainer for Deep Reinforcement Learning Policies
di: Samadi, Amir, et al.
Pubblicazione: (2024)
di: Samadi, Amir, et al.
Pubblicazione: (2024)
PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback
di: Chakraborty, Souradip, et al.
Pubblicazione: (2023)
di: Chakraborty, Souradip, et al.
Pubblicazione: (2023)
Mitigating the Alignment Tax of RLHF
di: Lin, Yong, et al.
Pubblicazione: (2023)
di: Lin, Yong, et al.
Pubblicazione: (2023)
HERO: Human-Feedback Efficient Reinforcement Learning for Online Diffusion Model Finetuning
di: Hiranaka, Ayano, et al.
Pubblicazione: (2024)
di: Hiranaka, Ayano, et al.
Pubblicazione: (2024)
Reinforcement Learning from Human Feedback
di: Lambert, Nathan
Pubblicazione: (2025)
di: Lambert, Nathan
Pubblicazione: (2025)
Trajectory Entropy Reinforcement Learning for Predictable and Robust Control
di: You, Bang, et al.
Pubblicazione: (2025)
di: You, Bang, et al.
Pubblicazione: (2025)
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
di: Xiong, Wei, et al.
Pubblicazione: (2023)
di: Xiong, Wei, et al.
Pubblicazione: (2023)
SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF
di: Chegini, Atoosa, et al.
Pubblicazione: (2024)
di: Chegini, Atoosa, et al.
Pubblicazione: (2024)
Strategyproof Reinforcement Learning from Human Feedback
di: Buening, Thomas Kleine, et al.
Pubblicazione: (2025)
di: Buening, Thomas Kleine, et al.
Pubblicazione: (2025)
MaxMin-RLHF: Alignment with Diverse Human Preferences
di: Chakraborty, Souradip, et al.
Pubblicazione: (2024)
di: Chakraborty, Souradip, et al.
Pubblicazione: (2024)
Failure Modes of Maximum Entropy RLHF
di: Çağatan, Ömer Veysel, et al.
Pubblicazione: (2025)
di: Çağatan, Ömer Veysel, et al.
Pubblicazione: (2025)
Explaining and Preventing Alignment Collapse in Iterative RLHF
di: Gauthier, Etienne, et al.
Pubblicazione: (2026)
di: Gauthier, Etienne, et al.
Pubblicazione: (2026)
Robust Reinforcement Learning from Corrupted Human Feedback
di: Bukharin, Alexander, et al.
Pubblicazione: (2024)
di: Bukharin, Alexander, et al.
Pubblicazione: (2024)
Dual Active Learning for Reinforcement Learning from Human Feedback
di: Liu, Pangpang, et al.
Pubblicazione: (2024)
di: Liu, Pangpang, et al.
Pubblicazione: (2024)
Evaluating Defences against Unsafe Feedback in RLHF
di: Rosati, Domenic, et al.
Pubblicazione: (2024)
di: Rosati, Domenic, et al.
Pubblicazione: (2024)
Understanding the Learning Dynamics of Alignment with Human Feedback
di: Im, Shawn, et al.
Pubblicazione: (2024)
di: Im, Shawn, et al.
Pubblicazione: (2024)
Flexible Blood Glucose Control: Offline Reinforcement Learning from Human Feedback
di: Emerson, Harry, et al.
Pubblicazione: (2025)
di: Emerson, Harry, et al.
Pubblicazione: (2025)
Beyond RLHF: A Unified Theoretical Framework of Alignment
di: Yun, Jihun, et al.
Pubblicazione: (2025)
di: Yun, Jihun, et al.
Pubblicazione: (2025)
Unifying Stable Optimization and Reference Regularization in RLHF
di: He, Li, et al.
Pubblicazione: (2026)
di: He, Li, et al.
Pubblicazione: (2026)
Enhancing RLHF with Human Gaze Modeling
di: Galliamov, Karim, et al.
Pubblicazione: (2025)
di: Galliamov, Karim, et al.
Pubblicazione: (2025)
Reinforcement Learning from Human Feedback: A Statistical Perspective
di: Liu, Pangpang, et al.
Pubblicazione: (2026)
di: Liu, Pangpang, et al.
Pubblicazione: (2026)
Reinforcement Learning from Multi-level and Episodic Human Feedback
di: Elahi, Muhammad Qasim, et al.
Pubblicazione: (2025)
di: Elahi, Muhammad Qasim, et al.
Pubblicazione: (2025)
A Minimaximalist Approach to Reinforcement Learning from Human Feedback
di: Swamy, Gokul, et al.
Pubblicazione: (2024)
di: Swamy, Gokul, et al.
Pubblicazione: (2024)
Dense Reward for Free in Reinforcement Learning from Human Feedback
di: Chan, Alex J., et al.
Pubblicazione: (2024)
di: Chan, Alex J., et al.
Pubblicazione: (2024)
Multi-turn Reinforcement Learning from Preference Human Feedback
di: Shani, Lior, et al.
Pubblicazione: (2024)
di: Shani, Lior, et al.
Pubblicazione: (2024)
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
di: Hahm, Dongyoon, et al.
Pubblicazione: (2026)
di: Hahm, Dongyoon, et al.
Pubblicazione: (2026)
RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation
di: Park, Chanwoo, et al.
Pubblicazione: (2024)
di: Park, Chanwoo, et al.
Pubblicazione: (2024)
Solving the Inverse Alignment Problem for Efficient RLHF
di: Krishna, Shambhavi, et al.
Pubblicazione: (2024)
di: Krishna, Shambhavi, et al.
Pubblicazione: (2024)
Towards Reliable Alignment: Uncertainty-aware RLHF
di: Banerjee, Debangshu, et al.
Pubblicazione: (2024)
di: Banerjee, Debangshu, et al.
Pubblicazione: (2024)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2025)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2025)
The Power of Active Multi-Task Learning in Reinforcement Learning from Human Feedback
di: Chen, Ruitao, et al.
Pubblicazione: (2024)
di: Chen, Ruitao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
AuON: A Linear-time Alternative to Orthogonal Momentum Updates
di: Maity, Dipan
Pubblicazione: (2025) -
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
di: Lee, Harrison, et al.
Pubblicazione: (2023) -
Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
di: Ji, Jiaming, et al.
Pubblicazione: (2025) -
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
di: Chaudhari, Shreyas, et al.
Pubblicazione: (2024) -
Gated-SwinRMT: Unifying Swin Windowed Attention with Retentive Manhattan Decay via Input-Dependent Gating
di: Maity, Dipan, et al.
Pubblicazione: (2026)