A Generalized Acquisition Function for Preference-based Reward Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Ellis, Evan, Ghosal, Gaurav R., Russell, Stuart J., Dragan, Anca, Bıyık, Erdem |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Batch Active Learning of Reward Functions from Human Preferences
di: Bıyık, Erdem, et al.
Pubblicazione: (2024)
di: Bıyık, Erdem, et al.
Pubblicazione: (2024)
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
di: Hwang, Minjune, et al.
Pubblicazione: (2026)
di: Hwang, Minjune, et al.
Pubblicazione: (2026)
MILE: Model-based Intervention Learning
di: Korkmaz, Yigit, et al.
Pubblicazione: (2025)
di: Korkmaz, Yigit, et al.
Pubblicazione: (2025)
Actor-Free Continuous Control via Structurally Maximizable Q-Functions
di: Korkmaz, Yigit, et al.
Pubblicazione: (2025)
di: Korkmaz, Yigit, et al.
Pubblicazione: (2025)
GABRIL: Gaze-Based Regularization for Mitigating Causal Confusion in Imitation Learning
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2025)
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2025)
AI Alignment with Changing and Influenceable Reward Functions
di: Carroll, Micah, et al.
Pubblicazione: (2024)
di: Carroll, Micah, et al.
Pubblicazione: (2024)
When a Robot is More Capable than a Human: Learning from Constrained Demonstrators
di: Li, Xinhu, et al.
Pubblicazione: (2025)
di: Li, Xinhu, et al.
Pubblicazione: (2025)
Coprocessor Actor Critic: A Model-Based Reinforcement Learning Approach For Adaptive Brain Stimulation
di: Pan, Michelle, et al.
Pubblicazione: (2024)
di: Pan, Michelle, et al.
Pubblicazione: (2024)
Trajectory Improvement and Reward Learning from Comparative Language Feedback
di: Yang, Zhaojing, et al.
Pubblicazione: (2024)
di: Yang, Zhaojing, et al.
Pubblicazione: (2024)
CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations
di: Liang, Anthony, et al.
Pubblicazione: (2025)
di: Liang, Anthony, et al.
Pubblicazione: (2025)
IMPACT: Intelligent Motion Planning with Acceptable Contact Trajectories via Vision-Language Models
di: Ling, Yiyang, et al.
Pubblicazione: (2025)
di: Ling, Yiyang, et al.
Pubblicazione: (2025)
Multi-Agent Inverse Q-Learning from Demonstrations
di: Haynam, Nathaniel, et al.
Pubblicazione: (2025)
di: Haynam, Nathaniel, et al.
Pubblicazione: (2025)
Learning to Assist Humans without Inferring Rewards
di: Myers, Vivek, et al.
Pubblicazione: (2024)
di: Myers, Vivek, et al.
Pubblicazione: (2024)
EXTRACT: Efficient Policy Learning by Extracting Transferable Robot Skills from Offline Data
di: Zhang, Jesse, et al.
Pubblicazione: (2024)
di: Zhang, Jesse, et al.
Pubblicazione: (2024)
Residual Reward Models for Preference-based Reinforcement Learning
di: Cao, Chenyang, et al.
Pubblicazione: (2025)
di: Cao, Chenyang, et al.
Pubblicazione: (2025)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
di: Wang, Yufei, et al.
Pubblicazione: (2024)
di: Wang, Yufei, et al.
Pubblicazione: (2024)
Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
di: Liang, Anthony, et al.
Pubblicazione: (2026)
di: Liang, Anthony, et al.
Pubblicazione: (2026)
Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions
di: Jain, Ayush, et al.
Pubblicazione: (2024)
di: Jain, Ayush, et al.
Pubblicazione: (2024)
PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies
di: Zhang, Jesse, et al.
Pubblicazione: (2025)
di: Zhang, Jesse, et al.
Pubblicazione: (2025)
Aligning Robot and Human Representations
di: Bobu, Andreea, et al.
Pubblicazione: (2023)
di: Bobu, Andreea, et al.
Pubblicazione: (2023)
The Effective Horizon Explains Deep RL Performance in Stochastic Environments
di: Laidlaw, Cassidy, et al.
Pubblicazione: (2023)
di: Laidlaw, Cassidy, et al.
Pubblicazione: (2023)
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
di: Ishihara, Yu, et al.
Pubblicazione: (2025)
di: Ishihara, Yu, et al.
Pubblicazione: (2025)
A Review of Reward Functions for Reinforcement Learning in the context of Autonomous Driving
di: Abouelazm, Ahmed, et al.
Pubblicazione: (2024)
di: Abouelazm, Ahmed, et al.
Pubblicazione: (2024)
Cross-Domain Imitation Learning via Optimal Transport
di: Fickinger, Arnaud, et al.
Pubblicazione: (2021)
di: Fickinger, Arnaud, et al.
Pubblicazione: (2021)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
di: Laidlaw, Cassidy, et al.
Pubblicazione: (2024)
di: Laidlaw, Cassidy, et al.
Pubblicazione: (2024)
RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences
di: Cheng, Jie, et al.
Pubblicazione: (2024)
di: Cheng, Jie, et al.
Pubblicazione: (2024)
ViSaRL: Visual Reinforcement Learning Guided by Human Saliency
di: Liang, Anthony, et al.
Pubblicazione: (2024)
di: Liang, Anthony, et al.
Pubblicazione: (2024)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
di: Lang, Leon, et al.
Pubblicazione: (2024)
di: Lang, Leon, et al.
Pubblicazione: (2024)
Training LLM Agents to Empower Humans
di: Ellis, Evan, et al.
Pubblicazione: (2025)
di: Ellis, Evan, et al.
Pubblicazione: (2025)
Zero-Shot Visual Generalization in Robot Manipulation
di: Batra, Sumeet, et al.
Pubblicazione: (2025)
di: Batra, Sumeet, et al.
Pubblicazione: (2025)
Improving User Experience in Preference-Based Optimization of Reward Functions for Assistive Robots
di: Dennler, Nathaniel, et al.
Pubblicazione: (2024)
di: Dennler, Nathaniel, et al.
Pubblicazione: (2024)
DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
di: Lee, Vint, et al.
Pubblicazione: (2023)
di: Lee, Vint, et al.
Pubblicazione: (2023)
Object and Relation Centric Representations for Push Effect Prediction
di: Tekden, Ahmet E., et al.
Pubblicazione: (2021)
di: Tekden, Ahmet E., et al.
Pubblicazione: (2021)
Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
di: Zhang, Rongtao, et al.
Pubblicazione: (2026)
di: Zhang, Rongtao, et al.
Pubblicazione: (2026)
Diffusion-Reward Adversarial Imitation Learning
di: Lai, Chun-Mao, et al.
Pubblicazione: (2024)
di: Lai, Chun-Mao, et al.
Pubblicazione: (2024)
Reward-Punishment Reinforcement Learning with Maximum Entropy
di: Wang, Jiexin, et al.
Pubblicazione: (2024)
di: Wang, Jiexin, et al.
Pubblicazione: (2024)
TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
di: Liu, Yuyang, et al.
Pubblicazione: (2025)
di: Liu, Yuyang, et al.
Pubblicazione: (2025)
Subwords as Skills: Tokenization for Sparse-Reward Reinforcement Learning
di: Yunis, David, et al.
Pubblicazione: (2023)
di: Yunis, David, et al.
Pubblicazione: (2023)
Adaptive Querying for Reward Learning from Human Feedback
di: Anand, Yashwanthi, et al.
Pubblicazione: (2024)
di: Anand, Yashwanthi, et al.
Pubblicazione: (2024)
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
di: Diaz-Bone, Leander, et al.
Pubblicazione: (2025)
di: Diaz-Bone, Leander, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Batch Active Learning of Reward Functions from Human Preferences
di: Bıyık, Erdem, et al.
Pubblicazione: (2024) -
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
di: Hwang, Minjune, et al.
Pubblicazione: (2026) -
MILE: Model-based Intervention Learning
di: Korkmaz, Yigit, et al.
Pubblicazione: (2025) -
Actor-Free Continuous Control via Structurally Maximizable Q-Functions
di: Korkmaz, Yigit, et al.
Pubblicazione: (2025) -
GABRIL: Gaze-Based Regularization for Mitigating Causal Confusion in Imitation Learning
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2025)