Batch Active Learning of Reward Functions from Human Preferences
Fuente:
arXiv
Saved in:
| Main Authors: | Bıyık, Erdem, Anari, Nima, Sadigh, Dorsa |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
by: Hwang, Minjune, et al.
Published: (2026)
by: Hwang, Minjune, et al.
Published: (2026)
A Generalized Acquisition Function for Preference-based Reward Learning
by: Ellis, Evan, et al.
Published: (2024)
by: Ellis, Evan, et al.
Published: (2024)
MILE: Model-based Intervention Learning
by: Korkmaz, Yigit, et al.
Published: (2025)
by: Korkmaz, Yigit, et al.
Published: (2025)
Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Autospeculation
by: Hu, Hengyuan, et al.
Published: (2025)
by: Hu, Hengyuan, et al.
Published: (2025)
When a Robot is More Capable than a Human: Learning from Constrained Demonstrators
by: Li, Xinhu, et al.
Published: (2025)
by: Li, Xinhu, et al.
Published: (2025)
Actor-Free Continuous Control via Structurally Maximizable Q-Functions
by: Korkmaz, Yigit, et al.
Published: (2025)
by: Korkmaz, Yigit, et al.
Published: (2025)
Motion Tracks: A Unified Representation for Human-Robot Transfer in Few-Shot Imitation Learning
by: Ren, Juntao, et al.
Published: (2025)
by: Ren, Juntao, et al.
Published: (2025)
GABRIL: Gaze-Based Regularization for Mitigating Causal Confusion in Imitation Learning
by: Banayeeanzade, Amin, et al.
Published: (2025)
by: Banayeeanzade, Amin, et al.
Published: (2025)
CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations
by: Liang, Anthony, et al.
Published: (2025)
by: Liang, Anthony, et al.
Published: (2025)
What Matters for Batch Online Reinforcement Learning in Robotics?
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
Multi-Agent Inverse Q-Learning from Demonstrations
by: Haynam, Nathaniel, et al.
Published: (2025)
by: Haynam, Nathaniel, et al.
Published: (2025)
Efficient Data Collection for Robotic Manipulation via Compositional Generalization
by: Gao, Jensen, et al.
Published: (2024)
by: Gao, Jensen, et al.
Published: (2024)
IMPACT: Intelligent Motion Planning with Acceptable Contact Trajectories via Vision-Language Models
by: Ling, Yiyang, et al.
Published: (2025)
by: Ling, Yiyang, et al.
Published: (2025)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
by: Wang, Yufei, et al.
Published: (2024)
by: Wang, Yufei, et al.
Published: (2024)
Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning
by: Clark, Jaden, et al.
Published: (2025)
by: Clark, Jaden, et al.
Published: (2025)
EXTRACT: Efficient Policy Learning by Extracting Transferable Robot Skills from Offline Data
by: Zhang, Jesse, et al.
Published: (2024)
by: Zhang, Jesse, et al.
Published: (2024)
Vocal Sandbox: Continual Learning and Adaptation for Situated Human-Robot Collaboration
by: Grannen, Jennifer, et al.
Published: (2024)
by: Grannen, Jennifer, et al.
Published: (2024)
ProVox: Personalization and Proactive Planning for Situated Human-Robot Collaboration
by: Grannen, Jennifer, et al.
Published: (2025)
by: Grannen, Jennifer, et al.
Published: (2025)
A Taxonomy for Evaluating Generalist Robot Manipulation Policies
by: Gao, Jensen, et al.
Published: (2025)
by: Gao, Jensen, et al.
Published: (2025)
Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions
by: Jain, Ayush, et al.
Published: (2024)
by: Jain, Ayush, et al.
Published: (2024)
What's the Move? Hybrid Imitation Learning via Salient Points
by: Sundaresan, Priya, et al.
Published: (2024)
by: Sundaresan, Priya, et al.
Published: (2024)
ViSaRL: Visual Reinforcement Learning Guided by Human Saliency
by: Liang, Anthony, et al.
Published: (2024)
by: Liang, Anthony, et al.
Published: (2024)
Residual Reward Models for Preference-based Reinforcement Learning
by: Cao, Chenyang, et al.
Published: (2025)
by: Cao, Chenyang, et al.
Published: (2025)
Contrastive Preference Learning: Learning from Human Feedback without RL
by: Hejna, Joey, et al.
Published: (2023)
by: Hejna, Joey, et al.
Published: (2023)
Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
by: Liang, Anthony, et al.
Published: (2026)
by: Liang, Anthony, et al.
Published: (2026)
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
by: Ishihara, Yu, et al.
Published: (2025)
by: Ishihara, Yu, et al.
Published: (2025)
Predictive Preference Learning from Human Interventions
by: Cai, Haoyuan, et al.
Published: (2025)
by: Cai, Haoyuan, et al.
Published: (2025)
Adaptive Querying for Reward Learning from Human Feedback
by: Anand, Yashwanthi, et al.
Published: (2024)
by: Anand, Yashwanthi, et al.
Published: (2024)
Distilling and Retrieving Generalizable Knowledge for Robot Manipulation via Language Corrections
by: Zha, Lihan, et al.
Published: (2023)
by: Zha, Lihan, et al.
Published: (2023)
PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies
by: Zhang, Jesse, et al.
Published: (2025)
by: Zhang, Jesse, et al.
Published: (2025)
Imitation Bootstrapped Reinforcement Learning
by: Hu, Hengyuan, et al.
Published: (2023)
by: Hu, Hengyuan, et al.
Published: (2023)
How to Train Your Robots? The Impact of Demonstration Modality on Imitation Learning
by: Li, Haozhuo, et al.
Published: (2025)
by: Li, Haozhuo, et al.
Published: (2025)
A Review of Reward Functions for Reinforcement Learning in the context of Autonomous Driving
by: Abouelazm, Ahmed, et al.
Published: (2024)
by: Abouelazm, Ahmed, et al.
Published: (2024)
Explore until Confident: Efficient Exploration for Embodied Question Answering
by: Ren, Allen Z., et al.
Published: (2024)
by: Ren, Allen Z., et al.
Published: (2024)
Latent Diffusion Planning for Imitation Learning
by: Xie, Amber, et al.
Published: (2025)
by: Xie, Amber, et al.
Published: (2025)
Data Retrieval with Importance Weights for Few-Shot Imitation Learning
by: Xie, Amber, et al.
Published: (2025)
by: Xie, Amber, et al.
Published: (2025)
Improving User Experience in Preference-Based Optimization of Reward Functions for Assistive Robots
by: Dennler, Nathaniel, et al.
Published: (2024)
by: Dennler, Nathaniel, et al.
Published: (2024)
Aligning Robot Navigation Behaviors with Human Intentions and Preferences
by: Karnan, Haresh
Published: (2024)
by: Karnan, Haresh
Published: (2024)
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
by: Poddar, Sriyash, et al.
Published: (2024)
by: Poddar, Sriyash, et al.
Published: (2024)
TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
by: Liu, Yuyang, et al.
Published: (2025)
by: Liu, Yuyang, et al.
Published: (2025)
Similar Items
-
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
by: Hwang, Minjune, et al.
Published: (2026) -
A Generalized Acquisition Function for Preference-based Reward Learning
by: Ellis, Evan, et al.
Published: (2024) -
MILE: Model-based Intervention Learning
by: Korkmaz, Yigit, et al.
Published: (2025) -
Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Autospeculation
by: Hu, Hengyuan, et al.
Published: (2025) -
When a Robot is More Capable than a Human: Learning from Constrained Demonstrators
by: Li, Xinhu, et al.
Published: (2025)