Hindsight PRIORs for Reward Learning from Human Preferences
Fuente:
arXiv
Salvato in:
| Autori principali: | Verma, Mudit, Metcalf, Katherine |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning a Canonical Basis of Human Preferences from Binary Ratings
di: Vodrahalli, Kailas, et al.
Pubblicazione: (2025)
di: Vodrahalli, Kailas, et al.
Pubblicazione: (2025)
Learning to Assist Humans without Inferring Rewards
di: Myers, Vivek, et al.
Pubblicazione: (2024)
di: Myers, Vivek, et al.
Pubblicazione: (2024)
Influencing Humans to Conform to Preference Models for RLHF
di: Hatgis-Kessell, Stephane, et al.
Pubblicazione: (2025)
di: Hatgis-Kessell, Stephane, et al.
Pubblicazione: (2025)
Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
di: Zhang, Rongtao, et al.
Pubblicazione: (2026)
di: Zhang, Rongtao, et al.
Pubblicazione: (2026)
Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?
di: Verma, Mudit, et al.
Pubblicazione: (2024)
di: Verma, Mudit, et al.
Pubblicazione: (2024)
Improving User Experience in Preference-Based Optimization of Reward Functions for Assistive Robots
di: Dennler, Nathaniel, et al.
Pubblicazione: (2024)
di: Dennler, Nathaniel, et al.
Pubblicazione: (2024)
PREDICT: Preference Reasoning by Evaluating Decomposed preferences Inferred from Candidate Trajectories
di: Aroca-Ouellette, Stephane, et al.
Pubblicazione: (2024)
di: Aroca-Ouellette, Stephane, et al.
Pubblicazione: (2024)
Towards Human Haptic Gesture Interpretation for Robotic Systems
di: Bianchini, Bibit, et al.
Pubblicazione: (2020)
di: Bianchini, Bibit, et al.
Pubblicazione: (2020)
Integrating Policy Summaries with Reward Decomposition for Explaining Reinforcement Learning Agents
di: Septon, Yael, et al.
Pubblicazione: (2022)
di: Septon, Yael, et al.
Pubblicazione: (2022)
Enhancing Preference-based Linear Bandits via Human Response Time
di: Li, Shen, et al.
Pubblicazione: (2024)
di: Li, Shen, et al.
Pubblicazione: (2024)
Optimizing Data Delivery: Insights from User Preferences on Visuals, Tables, and Text
di: Luera, Reuben, et al.
Pubblicazione: (2024)
di: Luera, Reuben, et al.
Pubblicazione: (2024)
Contrastive Learning from Exploratory Actions: Leveraging Natural Interactions for Preference Elicitation
di: Dennler, Nathaniel, et al.
Pubblicazione: (2025)
di: Dennler, Nathaniel, et al.
Pubblicazione: (2025)
Cost-Effective Online Multi-LLM Selection with Versatile Reward Models
di: Dai, Xiangxiang, et al.
Pubblicazione: (2024)
di: Dai, Xiangxiang, et al.
Pubblicazione: (2024)
Explanation through Reward Model Reconciliation using POMDP Tree Search
di: Kraske, Benjamin D., et al.
Pubblicazione: (2023)
di: Kraske, Benjamin D., et al.
Pubblicazione: (2023)
Learning to Decide with AI Assistance under Human-Alignment
di: Benz, Nina Corvelo, et al.
Pubblicazione: (2026)
di: Benz, Nina Corvelo, et al.
Pubblicazione: (2026)
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
di: Singh, Anukriti, et al.
Pubblicazione: (2025)
di: Singh, Anukriti, et al.
Pubblicazione: (2025)
Making RL with Preference-based Feedback Efficient via Randomization
di: Wu, Runzhe, et al.
Pubblicazione: (2023)
di: Wu, Runzhe, et al.
Pubblicazione: (2023)
Integrating Human Expertise in Continuous Spaces: A Novel Interactive Bayesian Optimization Framework with Preference Expected Improvement
di: Feith, Nikolaus, et al.
Pubblicazione: (2024)
di: Feith, Nikolaus, et al.
Pubblicazione: (2024)
AI, Meet Human: Learning Paradigms for Hybrid Decision Making Systems
di: Punzi, Clara, et al.
Pubblicazione: (2024)
di: Punzi, Clara, et al.
Pubblicazione: (2024)
MENTOR: Guiding Hierarchical Reinforcement Learning with Human Feedback and Dynamic Distance Constraint
di: Zhou, Xinglin, et al.
Pubblicazione: (2024)
di: Zhou, Xinglin, et al.
Pubblicazione: (2024)
Predictive AI Can Support Human Learning while Preserving Error Diversity
di: He, Vivianna Fang, et al.
Pubblicazione: (2025)
di: He, Vivianna Fang, et al.
Pubblicazione: (2025)
Co-Creative Learning via Metropolis-Hastings Interaction between Humans and AI
di: Okumura, Ryota, et al.
Pubblicazione: (2025)
di: Okumura, Ryota, et al.
Pubblicazione: (2025)
Efficient Human-in-the-Loop Active Learning: A Novel Framework for Data Labeling in AI Systems
di: Huang, Yiran, et al.
Pubblicazione: (2024)
di: Huang, Yiran, et al.
Pubblicazione: (2024)
On the Utility of Accounting for Human Beliefs about AI Intention in Human-AI Collaboration
di: Yu, Guanghui, et al.
Pubblicazione: (2024)
di: Yu, Guanghui, et al.
Pubblicazione: (2024)
Building Machines that Learn and Think with People
di: Collins, Katherine M., et al.
Pubblicazione: (2024)
di: Collins, Katherine M., et al.
Pubblicazione: (2024)
Human Expertise in Algorithmic Prediction
di: Alur, Rohan, et al.
Pubblicazione: (2024)
di: Alur, Rohan, et al.
Pubblicazione: (2024)
Human-Computer Interaction and Human-AI Collaboration in Advanced Air Mobility: A Comprehensive Review
di: Sagirli, Fatma Yamac, et al.
Pubblicazione: (2024)
di: Sagirli, Fatma Yamac, et al.
Pubblicazione: (2024)
Interaction Dynamics as a Reward Signal for LLMs
di: Gooding, Sian, et al.
Pubblicazione: (2025)
di: Gooding, Sian, et al.
Pubblicazione: (2025)
Align When They Want, Complement When They Need! Human-Centered Ensembles for Adaptive Human-AI Collaboration
di: Amin, Hasan, et al.
Pubblicazione: (2026)
di: Amin, Hasan, et al.
Pubblicazione: (2026)
HAPI: A Model for Learning Robot Facial Expressions from Human Preferences
di: Yang, Dongsheng, et al.
Pubblicazione: (2025)
di: Yang, Dongsheng, et al.
Pubblicazione: (2025)
Does Calibration Affect Human Actions?
di: Nizri, Meir, et al.
Pubblicazione: (2025)
di: Nizri, Meir, et al.
Pubblicazione: (2025)
Human-AI Collaborative Uncertainty Quantification
di: Noorani, Sima, et al.
Pubblicazione: (2025)
di: Noorani, Sima, et al.
Pubblicazione: (2025)
Trustworthy Human-AI Collaboration: Reinforcement Learning with Human Feedback and Physics Knowledge for Safe Autonomous Driving
di: Huang, Zilin, et al.
Pubblicazione: (2024)
di: Huang, Zilin, et al.
Pubblicazione: (2024)
CREW: Facilitating Human-AI Teaming Research
di: Zhang, Lingyu, et al.
Pubblicazione: (2024)
di: Zhang, Lingyu, et al.
Pubblicazione: (2024)
Reassessing Evaluation Functions in Algorithmic Recourse: An Empirical Study from a Human-Centered Perspective
di: Tominaga, Tomu, et al.
Pubblicazione: (2024)
di: Tominaga, Tomu, et al.
Pubblicazione: (2024)
BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback
di: Pandey, Gaurav, et al.
Pubblicazione: (2024)
di: Pandey, Gaurav, et al.
Pubblicazione: (2024)
Optimizing Delegation in Collaborative Human-AI Hybrid Teams
di: Fuchs, Andrew, et al.
Pubblicazione: (2024)
di: Fuchs, Andrew, et al.
Pubblicazione: (2024)
A No Free Lunch Theorem for Human-AI Collaboration
di: Peng, Kenny, et al.
Pubblicazione: (2024)
di: Peng, Kenny, et al.
Pubblicazione: (2024)
Strength Estimation and Human-Like Strength Adjustment in Games
di: Chen, Chun Jung, et al.
Pubblicazione: (2025)
di: Chen, Chun Jung, et al.
Pubblicazione: (2025)
Advancing Human-Machine Teaming: Concepts, Challenges, and Applications
di: Chen, Dian, et al.
Pubblicazione: (2025)
di: Chen, Dian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Learning a Canonical Basis of Human Preferences from Binary Ratings
di: Vodrahalli, Kailas, et al.
Pubblicazione: (2025) -
Learning to Assist Humans without Inferring Rewards
di: Myers, Vivek, et al.
Pubblicazione: (2024) -
Influencing Humans to Conform to Preference Models for RLHF
di: Hatgis-Kessell, Stephane, et al.
Pubblicazione: (2025) -
Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
di: Zhang, Rongtao, et al.
Pubblicazione: (2026) -
Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?
di: Verma, Mudit, et al.
Pubblicazione: (2024)