Governance Challenges in Reinforcement Learning from Human Feedback: Evaluator Rationality and Reinforcement Stability
Fuente:
arXiv
Saved in:
| Main Authors: | Alsagheer, Dana, Kamal, Abdulrahman, Kamal, Mohammad, Shi, Weidong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From General Reasoning to Domain Expertise: Uncovering the Limits of Generalization in Large Language Models
by: Alsagheer, Dana, et al.
Published: (2025)
by: Alsagheer, Dana, et al.
Published: (2025)
Detecting Conspiracy Theory Against COVID-19 Vaccines
by: Amin, Md Hasibul, et al.
Published: (2022)
by: Amin, Md Hasibul, et al.
Published: (2022)
Comparing Rationality Between Large Language Models and Humans: Insights and Open Questions
by: Alsagheer, Dana, et al.
Published: (2024)
by: Alsagheer, Dana, et al.
Published: (2024)
Ethics and Persuasion in Reinforcement Learning from Human Feedback: A Procedural Rhetorical Approach
by: Lodoen, Shannon, et al.
Published: (2025)
by: Lodoen, Shannon, et al.
Published: (2025)
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
by: Peng, Xiyue, et al.
Published: (2024)
by: Peng, Xiyue, et al.
Published: (2024)
Reinforcement Learning from Human Feedback: Whose Culture, Whose Values, Whose Perspectives?
by: Barman, Kristian González, et al.
Published: (2024)
by: Barman, Kristian González, et al.
Published: (2024)
LLM Nepotism in Organizational Governance
by: Mao, Shunqi, et al.
Published: (2026)
by: Mao, Shunqi, et al.
Published: (2026)
Challenging the Machine: Contestability in Government AI Systems
by: Landau, Susan, et al.
Published: (2024)
by: Landau, Susan, et al.
Published: (2024)
AI Governance Control Stack for Operational Stability: Achieving Hardened Governance in AI Systems
by: Morgan, Horatio
Published: (2026)
by: Morgan, Horatio
Published: (2026)
Sparks of Rationality: Do Reasoning LLMs Align with Human Judgment and Choice?
by: Tak, Ala N., et al.
Published: (2026)
by: Tak, Ala N., et al.
Published: (2026)
FACTER: Fairness-Aware Conformal Thresholding and Prompt Engineering for Enabling Fair LLM-Based Recommender Systems
by: Fayyazi, Arya, et al.
Published: (2025)
by: Fayyazi, Arya, et al.
Published: (2025)
Responsible AI Governance: A Response to UN Interim Report on Governing AI for Humanity
by: Kiden, Sarah, et al.
Published: (2024)
by: Kiden, Sarah, et al.
Published: (2024)
Challenges and Best Practices in Corporate AI Governance:Lessons from the Biopharmaceutical Industry
by: Mökander, Jakob, et al.
Published: (2024)
by: Mökander, Jakob, et al.
Published: (2024)
Evolutionary Reinforcement Learning based AI tutor for Socratic Interdisciplinary Instruction
by: Jiang, Mei, et al.
Published: (2025)
by: Jiang, Mei, et al.
Published: (2025)
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
by: Wang, Peisong, et al.
Published: (2025)
by: Wang, Peisong, et al.
Published: (2025)
Aryabhata 2: Scaling Reinforcement Learning for Advanced STEM Reasoning
by: Rastogi, Ritvik, et al.
Published: (2026)
by: Rastogi, Ritvik, et al.
Published: (2026)
Evaluating Language Models for Generating and Judging Programming Feedback
by: Koutcheme, Charles, et al.
Published: (2024)
by: Koutcheme, Charles, et al.
Published: (2024)
Failing on Bias Mitigation: A Case Study on the Challenges of Fairness in Government Data
by: Bo, Hongbo, et al.
Published: (2026)
by: Bo, Hongbo, et al.
Published: (2026)
RLCP: A Reinforcement Learning-based Copyright Protection Method for Text-to-Image Diffusion Model
by: Shi, Zhuan, et al.
Published: (2024)
by: Shi, Zhuan, et al.
Published: (2024)
Reinforcement Learning and Life Cycle Assessment for a Circular Economy -- Towards Progressive Computer Science
by: Buchner, Johannes
Published: (2025)
by: Buchner, Johannes
Published: (2025)
Science Out of Its Ivory Tower: Improving Accessibility with Reinforcement Learning
by: Wang, Haining, et al.
Published: (2024)
by: Wang, Haining, et al.
Published: (2024)
Intelli-Planner: Towards Customized Urban Planning via Large Language Model Empowered Reinforcement Learning
by: Yong, Xixian, et al.
Published: (2026)
by: Yong, Xixian, et al.
Published: (2026)
Building Capacity for Artificial Intelligence in Africa: A Cross-Country Survey of Challenges and Governance Pathways
by: Aryee, Jeffrey N. A., et al.
Published: (2025)
by: Aryee, Jeffrey N. A., et al.
Published: (2025)
Examining the Challenges of Intellectual Property in AI-Generated Productions
by: Mazhar, Ali, et al.
Published: (2026)
by: Mazhar, Ali, et al.
Published: (2026)
Datacenters in the Desert: Feasibility and Sustainability of LLM Inference in the Middle East
by: Hassan, Lara, et al.
Published: (2025)
by: Hassan, Lara, et al.
Published: (2025)
H2-MARL: Multi-Agent Reinforcement Learning for Pareto Optimality in Hospital Capacity Strain and Human Mobility during Epidemic
by: Luo, Xueting, et al.
Published: (2025)
by: Luo, Xueting, et al.
Published: (2025)
Escaping the Agreement Trap: Defensibility Signals for Evaluating Rule-Governed AI
by: O'Herlihy, Michael, et al.
Published: (2026)
by: O'Herlihy, Michael, et al.
Published: (2026)
Online Advertisements with LLMs: Opportunities and Challenges
by: Feizi, Soheil, et al.
Published: (2023)
by: Feizi, Soheil, et al.
Published: (2023)
Cultural Rights and the Rights to Development in the Age of AI: Implications for Global Human Rights Governance
by: Kriebitz, Alexander, et al.
Published: (2025)
by: Kriebitz, Alexander, et al.
Published: (2025)
Culturally-Attuned Moral Machines: Implicit Learning of Human Value Systems by AI through Inverse Reinforcement Learning
by: Oliveira, Nigini, et al.
Published: (2023)
by: Oliveira, Nigini, et al.
Published: (2023)
EduGym: An Environment and Notebook Suite for Reinforcement Learning Education
by: Moerland, Thomas M., et al.
Published: (2023)
by: Moerland, Thomas M., et al.
Published: (2023)
Monitoring Fidelity of Online Reinforcement Learning Algorithms in Clinical Trials
by: Trella, Anna L., et al.
Published: (2024)
by: Trella, Anna L., et al.
Published: (2024)
A Comparative Study of Technical Writing Feedback Quality: Evaluating LLMs, SLMs, and Humans in Computer Science Topics
by: Liu, Suqing, et al.
Published: (2025)
by: Liu, Suqing, et al.
Published: (2025)
A Framework for Human-AI Q-Matrix Refinement: A NeuralCDM Evaluation
by: Zhang, Ying, et al.
Published: (2026)
by: Zhang, Ying, et al.
Published: (2026)
Opportunities and Challenges of Frontier Data Governance With Synthetic Data
by: Thakur, Madhavendra, et al.
Published: (2025)
by: Thakur, Madhavendra, et al.
Published: (2025)
Challenges for AI in Multimodal STEM Assessments: a Human-AI Comparison
by: de Chillaz, Aymeric, et al.
Published: (2025)
by: de Chillaz, Aymeric, et al.
Published: (2025)
Parameter Efficient Reinforcement Learning from Human Feedback
by: Sidahmed, Hakim, et al.
Published: (2024)
by: Sidahmed, Hakim, et al.
Published: (2024)
Integrating Reason-Based Moral Decision-Making in the Reinforcement Learning Architecture
by: Dargasz, Lisa
Published: (2025)
by: Dargasz, Lisa
Published: (2025)
Optimizing HIV Patient Engagement with Reinforcement Learning in Resource-Limited Settings
by: Periáñez, África, et al.
Published: (2024)
by: Periáñez, África, et al.
Published: (2024)
Empowering Epidemic Response: The Role of Reinforcement Learning in Infectious Disease Control
by: Liu, Mutong, et al.
Published: (2026)
by: Liu, Mutong, et al.
Published: (2026)
Similar Items
-
From General Reasoning to Domain Expertise: Uncovering the Limits of Generalization in Large Language Models
by: Alsagheer, Dana, et al.
Published: (2025) -
Detecting Conspiracy Theory Against COVID-19 Vaccines
by: Amin, Md Hasibul, et al.
Published: (2022) -
Comparing Rationality Between Large Language Models and Humans: Insights and Open Questions
by: Alsagheer, Dana, et al.
Published: (2024) -
Ethics and Persuasion in Reinforcement Learning from Human Feedback: A Procedural Rhetorical Approach
by: Lodoen, Shannon, et al.
Published: (2025) -
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
by: Peng, Xiyue, et al.
Published: (2024)