Adaptive Querying for Reward Learning from Human Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Anand, Yashwanthi, Nwagwu, Nnamdi, Sabbe, Kevin, Fitter, Naomi T., Saisubramanian, Sandhya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Objective Planning with Contextual Lexicographic Reward Preferences
by: Rustagi, Pulkit, et al.
Published: (2025)
by: Rustagi, Pulkit, et al.
Published: (2025)
Learning with Expert Abstractions for Efficient Multi-Task Continuous Control
by: Jewett, Jeff, et al.
Published: (2025)
by: Jewett, Jeff, et al.
Published: (2025)
Calibrating Biophysical Models for Grape Phenology Prediction via Multi-Task Learning
by: Solow, William, et al.
Published: (2025)
by: Solow, William, et al.
Published: (2025)
Uncovering Systemic and Environment Errors in Autonomous Systems Using Differential Testing
by: Anand, Yashwanthi, et al.
Published: (2025)
by: Anand, Yashwanthi, et al.
Published: (2025)
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
by: Hwang, Minjune, et al.
Published: (2026)
by: Hwang, Minjune, et al.
Published: (2026)
Batch Active Learning of Reward Functions from Human Preferences
by: Bıyık, Erdem, et al.
Published: (2024)
by: Bıyık, Erdem, et al.
Published: (2024)
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
by: Ishihara, Yu, et al.
Published: (2025)
by: Ishihara, Yu, et al.
Published: (2025)
TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
by: Liu, Yuyang, et al.
Published: (2025)
by: Liu, Yuyang, et al.
Published: (2025)
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
by: Poddar, Sriyash, et al.
Published: (2024)
by: Poddar, Sriyash, et al.
Published: (2024)
Diffusion-Reward Adversarial Imitation Learning
by: Lai, Chun-Mao, et al.
Published: (2024)
by: Lai, Chun-Mao, et al.
Published: (2024)
Reward Learning from Suboptimal Demonstrations with Applications in Surgical Electrocautery
by: Karimi, Zohre, et al.
Published: (2024)
by: Karimi, Zohre, et al.
Published: (2024)
Eureka: Human-Level Reward Design via Coding Large Language Models
by: Ma, Yecheng Jason, et al.
Published: (2023)
by: Ma, Yecheng Jason, et al.
Published: (2023)
Learning Transferable Latent User Preferences for Human-Aligned Decision Making
by: Hyk, Alina, et al.
Published: (2026)
by: Hyk, Alina, et al.
Published: (2026)
Reward-Punishment Reinforcement Learning with Maximum Entropy
by: Wang, Jiexin, et al.
Published: (2024)
by: Wang, Jiexin, et al.
Published: (2024)
Robot Policy Learning with Temporal Optimal Transport Reward
by: Fu, Yuwei, et al.
Published: (2024)
by: Fu, Yuwei, et al.
Published: (2024)
Subwords as Skills: Tokenization for Sparse-Reward Reinforcement Learning
by: Yunis, David, et al.
Published: (2023)
by: Yunis, David, et al.
Published: (2023)
Residual Reward Models for Preference-based Reinforcement Learning
by: Cao, Chenyang, et al.
Published: (2025)
by: Cao, Chenyang, et al.
Published: (2025)
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
by: Diaz-Bone, Leander, et al.
Published: (2025)
by: Diaz-Bone, Leander, et al.
Published: (2025)
Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
by: Zheng, Qinqing, et al.
Published: (2024)
by: Zheng, Qinqing, et al.
Published: (2024)
DrS: Learning Reusable Dense Rewards for Multi-Stage Tasks
by: Mu, Tongzhou, et al.
Published: (2024)
by: Mu, Tongzhou, et al.
Published: (2024)
A Generalized Acquisition Function for Preference-based Reward Learning
by: Ellis, Evan, et al.
Published: (2024)
by: Ellis, Evan, et al.
Published: (2024)
Enabling Option Learning in Sparse Rewards with Hindsight Experience Replay
by: Romio, Gabriel, et al.
Published: (2026)
by: Romio, Gabriel, et al.
Published: (2026)
CaRL: Learning Scalable Planning Policies with Simple Rewards
by: Jaeger, Bernhard, et al.
Published: (2025)
by: Jaeger, Bernhard, et al.
Published: (2025)
Long-term Safe Reinforcement Learning with Binary Feedback
by: Wachi, Akifumi, et al.
Published: (2024)
by: Wachi, Akifumi, et al.
Published: (2024)
Real-World Offline Reinforcement Learning from Vision Language Model Feedback
by: Venkataraman, Sreyas, et al.
Published: (2024)
by: Venkataraman, Sreyas, et al.
Published: (2024)
How Much Progress Did I Make? An Unexplored Human Feedback Signal for Teaching Robots
by: Yu, Hang, et al.
Published: (2024)
by: Yu, Hang, et al.
Published: (2024)
A Review of Reward Functions for Reinforcement Learning in the context of Autonomous Driving
by: Abouelazm, Ahmed, et al.
Published: (2024)
by: Abouelazm, Ahmed, et al.
Published: (2024)
Can We Really Learn One Representation to Optimize All Rewards?
by: Zheng, Chongyi, et al.
Published: (2026)
by: Zheng, Chongyi, et al.
Published: (2026)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
by: Zhang, Chen Bo Calvin, et al.
Published: (2024)
by: Zhang, Chen Bo Calvin, et al.
Published: (2024)
From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models
by: Gumbsch, Christian, et al.
Published: (2026)
by: Gumbsch, Christian, et al.
Published: (2026)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
by: Wang, Yufei, et al.
Published: (2024)
by: Wang, Yufei, et al.
Published: (2024)
Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented Imitation
by: Guo, Yihong, et al.
Published: (2024)
by: Guo, Yihong, et al.
Published: (2024)
Learning to Recover: Dynamic Reward Shaping with Wheel-Leg Coordination for Fallen Robots
by: Deng, Boyuan, et al.
Published: (2025)
by: Deng, Boyuan, et al.
Published: (2025)
DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
by: Lee, Vint, et al.
Published: (2023)
by: Lee, Vint, et al.
Published: (2023)
Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation
by: Patel, Bhrij, et al.
Published: (2023)
by: Patel, Bhrij, et al.
Published: (2023)
Sample-Efficient Expert Query Control in Active Imitation Learning via Conformal Prediction
by: Firouzkouhi, Arad, et al.
Published: (2025)
by: Firouzkouhi, Arad, et al.
Published: (2025)
Learning Adaptive Dexterous Grasping from Single Demonstrations
by: Shi, Liangzhi, et al.
Published: (2025)
by: Shi, Liangzhi, et al.
Published: (2025)
A Hybrid Modeling Framework for Crop Prediction Tasks via Dynamic Parameter Calibration and Multi-Task Learning
by: Solow, William, et al.
Published: (2026)
by: Solow, William, et al.
Published: (2026)
Predictive Preference Learning from Human Interventions
by: Cai, Haoyuan, et al.
Published: (2025)
by: Cai, Haoyuan, et al.
Published: (2025)
Error-Feedback Model for Output Correction in Bilateral Control-Based Imitation Learning
by: Sato, Hiroshi, et al.
Published: (2024)
by: Sato, Hiroshi, et al.
Published: (2024)
Similar Items
-
Multi-Objective Planning with Contextual Lexicographic Reward Preferences
by: Rustagi, Pulkit, et al.
Published: (2025) -
Learning with Expert Abstractions for Efficient Multi-Task Continuous Control
by: Jewett, Jeff, et al.
Published: (2025) -
Calibrating Biophysical Models for Grape Phenology Prediction via Multi-Task Learning
by: Solow, William, et al.
Published: (2025) -
Uncovering Systemic and Environment Errors in Autonomous Systems Using Differential Testing
by: Anand, Yashwanthi, et al.
Published: (2025) -
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
by: Hwang, Minjune, et al.
Published: (2026)