Influencing Humans to Conform to Preference Models for RLHF
Fuente:
arXiv
Saved in:
| Main Authors: | Hatgis-Kessell, Stephane, Knox, W. Bradley, Booth, Serena, Stone, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
by: Hatgis-Kessell, Stephane, et al.
Published: (2026)
by: Hatgis-Kessell, Stephane, et al.
Published: (2026)
Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback
by: Yuan, Yifu, et al.
Published: (2024)
by: Yuan, Yifu, et al.
Published: (2024)
State Your Intention to Steer Your Attention: An AI Assistant for Intentional Digital Living
by: Choi, Juheon, et al.
Published: (2025)
by: Choi, Juheon, et al.
Published: (2025)
Hindsight PRIORs for Reward Learning from Human Preferences
by: Verma, Mudit, et al.
Published: (2024)
by: Verma, Mudit, et al.
Published: (2024)
Learning a Canonical Basis of Human Preferences from Binary Ratings
by: Vodrahalli, Kailas, et al.
Published: (2025)
by: Vodrahalli, Kailas, et al.
Published: (2025)
Understanding Impact of Human Feedback via Influence Functions
by: Min, Taywon, et al.
Published: (2025)
by: Min, Taywon, et al.
Published: (2025)
Can Interpretability Layouts Influence Human Perception of Offensive Sentences?
by: Santos, Thiago Freitas dos, et al.
Published: (2024)
by: Santos, Thiago Freitas dos, et al.
Published: (2024)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
by: Zhang, Rongtao, et al.
Published: (2026)
by: Zhang, Rongtao, et al.
Published: (2026)
Enhancing Preference-based Linear Bandits via Human Response Time
by: Li, Shen, et al.
Published: (2024)
by: Li, Shen, et al.
Published: (2024)
Conformal Set-based Human-AI Complementarity with Multiple Experts
by: Paat, Helbert, et al.
Published: (2025)
by: Paat, Helbert, et al.
Published: (2025)
Making RL with Preference-based Feedback Efficient via Randomization
by: Wu, Runzhe, et al.
Published: (2023)
by: Wu, Runzhe, et al.
Published: (2023)
The Human-Data-Model Interaction Canvas for Visual Analytics
by: Bernard, Jürgen
Published: (2025)
by: Bernard, Jürgen
Published: (2025)
Integrating Human Expertise in Continuous Spaces: A Novel Interactive Bayesian Optimization Framework with Preference Expected Improvement
by: Feith, Nikolaus, et al.
Published: (2024)
by: Feith, Nikolaus, et al.
Published: (2024)
Optimizing Data Delivery: Insights from User Preferences on Visuals, Tables, and Text
by: Luera, Reuben, et al.
Published: (2024)
by: Luera, Reuben, et al.
Published: (2024)
ABScribe: Rapid Exploration & Organization of Multiple Writing Variations in Human-AI Co-Writing Tasks using Large Language Models
by: Reza, Mohi, et al.
Published: (2023)
by: Reza, Mohi, et al.
Published: (2023)
Evaluation of Human-Understandability of Global Model Explanations using Decision Tree
by: Sivaprasad, Adarsa, et al.
Published: (2023)
by: Sivaprasad, Adarsa, et al.
Published: (2023)
The Model Mastery Lifecycle: A Framework for Designing Human-AI Interaction
by: Chignell, Mark, et al.
Published: (2024)
by: Chignell, Mark, et al.
Published: (2024)
RuleEdit: Failure-Guided Human-AI Model Editing with Prospective Impact Preview
by: Lee, Min Hun, et al.
Published: (2026)
by: Lee, Min Hun, et al.
Published: (2026)
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
by: Singh, Anukriti, et al.
Published: (2025)
by: Singh, Anukriti, et al.
Published: (2025)
On the Utility of Accounting for Human Beliefs about AI Intention in Human-AI Collaboration
by: Yu, Guanghui, et al.
Published: (2024)
by: Yu, Guanghui, et al.
Published: (2024)
Harmful Traits of AI Companions
by: Knox, W. Bradley, et al.
Published: (2025)
by: Knox, W. Bradley, et al.
Published: (2025)
Human Expertise in Algorithmic Prediction
by: Alur, Rohan, et al.
Published: (2024)
by: Alur, Rohan, et al.
Published: (2024)
Human-Computer Interaction and Human-AI Collaboration in Advanced Air Mobility: A Comprehensive Review
by: Sagirli, Fatma Yamac, et al.
Published: (2024)
by: Sagirli, Fatma Yamac, et al.
Published: (2024)
Generalizable Error Modeling for Human Data Annotation: Evidence From an Industry-Scale Search Data Annotation Program
by: Peters, Heinrich, et al.
Published: (2023)
by: Peters, Heinrich, et al.
Published: (2023)
Align When They Want, Complement When They Need! Human-Centered Ensembles for Adaptive Human-AI Collaboration
by: Amin, Hasan, et al.
Published: (2026)
by: Amin, Hasan, et al.
Published: (2026)
Does Calibration Affect Human Actions?
by: Nizri, Meir, et al.
Published: (2025)
by: Nizri, Meir, et al.
Published: (2025)
Human-AI Collaborative Uncertainty Quantification
by: Noorani, Sima, et al.
Published: (2025)
by: Noorani, Sima, et al.
Published: (2025)
Conformalized Interactive Imitation Learning: Handling Expert Shift and Intermittent Feedback
by: Zhao, Michelle, et al.
Published: (2024)
by: Zhao, Michelle, et al.
Published: (2024)
CREW: Facilitating Human-AI Teaming Research
by: Zhang, Lingyu, et al.
Published: (2024)
by: Zhang, Lingyu, et al.
Published: (2024)
Strength Estimation and Human-Like Strength Adjustment in Games
by: Chen, Chun Jung, et al.
Published: (2025)
by: Chen, Chun Jung, et al.
Published: (2025)
Advancing Human-Machine Teaming: Concepts, Challenges, and Applications
by: Chen, Dian, et al.
Published: (2025)
by: Chen, Dian, et al.
Published: (2025)
MedSyn: Enhancing Diagnostics with Human-AI Collaboration
by: Sayin, Burcu, et al.
Published: (2025)
by: Sayin, Burcu, et al.
Published: (2025)
Optimizing Delegation in Collaborative Human-AI Hybrid Teams
by: Fuchs, Andrew, et al.
Published: (2024)
by: Fuchs, Andrew, et al.
Published: (2024)
AI Agents for Inventory Control: Human-LLM-OR Complementarity
by: Baek, Jackie, et al.
Published: (2026)
by: Baek, Jackie, et al.
Published: (2026)
Learning to Decide with AI Assistance under Human-Alignment
by: Benz, Nina Corvelo, et al.
Published: (2026)
by: Benz, Nina Corvelo, et al.
Published: (2026)
Rationalize: Shared Semantic Reasoning for Human-AI Alignment
by: Dasgupta, Aritra, et al.
Published: (2026)
by: Dasgupta, Aritra, et al.
Published: (2026)
Toward Human-AI Complementarity Across Diverse Tasks
by: Xu, Yuzheng, et al.
Published: (2026)
by: Xu, Yuzheng, et al.
Published: (2026)
A No Free Lunch Theorem for Human-AI Collaboration
by: Peng, Kenny, et al.
Published: (2024)
by: Peng, Kenny, et al.
Published: (2024)
ADEPTS: A Capability Framework for Human-Centered Agent Design
by: D'Oro, Pierluca, et al.
Published: (2025)
by: D'Oro, Pierluca, et al.
Published: (2025)
Similar Items
-
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
by: Hatgis-Kessell, Stephane, et al.
Published: (2026) -
Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback
by: Yuan, Yifu, et al.
Published: (2024) -
State Your Intention to Steer Your Attention: An AI Assistant for Intentional Digital Living
by: Choi, Juheon, et al.
Published: (2025) -
Hindsight PRIORs for Reward Learning from Human Preferences
by: Verma, Mudit, et al.
Published: (2024) -
Learning a Canonical Basis of Human Preferences from Binary Ratings
by: Vodrahalli, Kailas, et al.
Published: (2025)