RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chaudhari, Shreyas, Aggarwal, Pranjal, Murahari, Vishvak, Rajpurohit, Tanmay, Kalyan, Ashwin, Narasimhan, Karthik, Deshpande, Ameet, da Silva, Bruno Castro |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
GEO: Generative Engine Optimization
par: Aggarwal, Pranjal, et autres
Publié: (2023)
par: Aggarwal, Pranjal, et autres
Publié: (2023)
Probing AI Safety with Source Code
par: Narayan, Ujwal, et autres
Publié: (2025)
par: Narayan, Ujwal, et autres
Publié: (2025)
Agent Context Protocols Enhance Collective Inference
par: Bhardwaj, Devansh, et autres
Publié: (2025)
par: Bhardwaj, Devansh, et autres
Publié: (2025)
PersonaGym: Evaluating Persona Agents and LLMs
par: Samuel, Vinay, et autres
Publié: (2024)
par: Samuel, Vinay, et autres
Publié: (2024)
QualEval: Qualitative Evaluation for Model Improvement
par: Murahari, Vishvak, et autres
Publié: (2023)
par: Murahari, Vishvak, et autres
Publié: (2023)
Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation
par: Chaudhari, Shreyas, et autres
Publié: (2024)
par: Chaudhari, Shreyas, et autres
Publié: (2024)
Language Models can Subtly Deceive Without Lying: A Case Study on Strategic Phrasing in Legislation
par: Dogra, Atharvan, et autres
Publié: (2024)
par: Dogra, Atharvan, et autres
Publié: (2024)
Which Rewards Matter? Reward Selection for Reinforcement Learning under Limited Feedback
par: Chaudhari, Shreyas, et autres
Publié: (2025)
par: Chaudhari, Shreyas, et autres
Publié: (2025)
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
par: Gupta, Shashank, et autres
Publié: (2023)
par: Gupta, Shashank, et autres
Publié: (2023)
Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
par: Dogra, Atharvan, et autres
Publié: (2025)
par: Dogra, Atharvan, et autres
Publié: (2025)
LLMs are Superior Feedback Providers: Bootstrapping Reasoning for Lie Detection with Self-Generated Feedback
par: Banerjee, Tanushree, et autres
Publié: (2024)
par: Banerjee, Tanushree, et autres
Publié: (2024)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
par: Aggarwal, Pranjal, et autres
Publié: (2025)
par: Aggarwal, Pranjal, et autres
Publié: (2025)
Dimension-Free Parameterized Approximation Schemes for Hybrid Clustering
par: Gadekar, Ameet, et autres
Publié: (2025)
par: Gadekar, Ameet, et autres
Publié: (2025)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
par: Lee, Harrison, et autres
Publié: (2023)
par: Lee, Harrison, et autres
Publié: (2023)
MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions
par: Chai, Yekun, et autres
Publié: (2024)
par: Chai, Yekun, et autres
Publié: (2024)
Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
par: Ji, Jiaming, et autres
Publié: (2025)
par: Ji, Jiaming, et autres
Publié: (2025)
Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback
par: Yuan, Yifu, et autres
Publié: (2024)
par: Yuan, Yifu, et autres
Publié: (2024)
Clustering under Constraints: Efficient Parameterized Approximation Schemes
par: Bhore, Sujoy, et autres
Publié: (2025)
par: Bhore, Sujoy, et autres
Publié: (2025)
Programming with Pixels: Can Computer-Use Agents do Software Engineering?
par: Aggarwal, Pranjal, et autres
Publié: (2025)
par: Aggarwal, Pranjal, et autres
Publié: (2025)
Enhancing LLMs for Physics Problem-Solving using Reinforcement Learning with Human-AI Feedback
par: Anand, Avinash, et autres
Publié: (2024)
par: Anand, Avinash, et autres
Publié: (2024)
Vanishing sheaves and the geometric Whittaker model
par: Roman Bezrukavnikov, et autres
Publié: (2025)
par: Roman Bezrukavnikov, et autres
Publié: (2025)
Character Sheaves on Tori over Local Fields
par: Deshpande, Tanmay, et autres
Publié: (2023)
par: Deshpande, Tanmay, et autres
Publié: (2023)
Vanishing sheaves and the geometric Whittaker model
par: Bezrukavnikov, Roman, et autres
Publié: (2023)
par: Bezrukavnikov, Roman, et autres
Publié: (2023)
Minimum Envy Graphical House Allocation Beyond Identical Valuations
par: Inamdar, Tanmay, et autres
Publié: (2026)
par: Inamdar, Tanmay, et autres
Publié: (2026)
ACE-RLHF: Automated Code Evaluation and Socratic Feedback Generation Tool using Large Language Models and Reinforcement Learning with Human Feedback
par: Rahman, Tasnia, et autres
Publié: (2025)
par: Rahman, Tasnia, et autres
Publié: (2025)
SAFE: Stable Alignment Finetuning with Entropy-Aware Predictive Control for Reinforcement Learning from Human Feedback (RLHF)
par: Maity, Dipan
Publié: (2026)
par: Maity, Dipan
Publié: (2026)
Deciphering the origins and growth of supermassive black holes
par: Aggarwal, Yash
Publié: (2021)
par: Aggarwal, Yash
Publié: (2021)
RLHF Fine-Tuning of LLMs for Alignment with Implicit User Feedback in Conversational Recommenders
par: Yang, Zhongheng, et autres
Publié: (2025)
par: Yang, Zhongheng, et autres
Publié: (2025)
A Construction of the Symmetric Monoidal Structure of the Geometric Whittaker Model
par: Choudhury, Ashutosh Roy, et autres
Publié: (2024)
par: Choudhury, Ashutosh Roy, et autres
Publié: (2024)
ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback
par: Hou, Zhenyu, et autres
Publié: (2024)
par: Hou, Zhenyu, et autres
Publié: (2024)
A Survey of Reinforcement Learning For Economics
par: Rawat, Pranjal
Publié: (2026)
par: Rawat, Pranjal
Publié: (2026)
Approximating Auction Equilibria with Reinforcement Learning
par: Rawat, Pranjal
Publié: (2024)
par: Rawat, Pranjal
Publié: (2024)
When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback
par: Pan, Jane, et autres
Publié: (2025)
par: Pan, Jane, et autres
Publié: (2025)
A Critical Study of What Code-LLMs (Do Not) Learn
par: Anand, Abhinav, et autres
Publié: (2024)
par: Anand, Abhinav, et autres
Publié: (2024)
On The Global Convergence Of Online RLHF With Neural Parametrization
par: Gaur, Mudit, et autres
Publié: (2024)
par: Gaur, Mudit, et autres
Publié: (2024)
ReLU Networks as Random Functions: Their Distribution in Probability Space
par: Chaudhari, Shreyas, et autres
Publié: (2025)
par: Chaudhari, Shreyas, et autres
Publié: (2025)
AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and Treefinement
par: Aggarwal, Pranjal, et autres
Publié: (2024)
par: Aggarwal, Pranjal, et autres
Publié: (2024)
Gym-Anything: Turn any Software into an Agent Environment
par: Aggarwal, Pranjal, et autres
Publié: (2026)
par: Aggarwal, Pranjal, et autres
Publié: (2026)
Reward-Robust RLHF in LLMs
par: Yan, Yuzi, et autres
Publié: (2024)
par: Yan, Yuzi, et autres
Publié: (2024)
Evaluating Defences against Unsafe Feedback in RLHF
par: Rosati, Domenic, et autres
Publié: (2024)
par: Rosati, Domenic, et autres
Publié: (2024)
Documents similaires
-
GEO: Generative Engine Optimization
par: Aggarwal, Pranjal, et autres
Publié: (2023) -
Probing AI Safety with Source Code
par: Narayan, Ujwal, et autres
Publié: (2025) -
Agent Context Protocols Enhance Collective Inference
par: Bhardwaj, Devansh, et autres
Publié: (2025) -
PersonaGym: Evaluating Persona Agents and LLMs
par: Samuel, Vinay, et autres
Publié: (2024) -
QualEval: Qualitative Evaluation for Model Improvement
par: Murahari, Vishvak, et autres
Publié: (2023)