Causally Robust Reward Learning from Reason-Augmented Preference Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Hwang, Minjune, Korkmaz, Yigit, Seita, Daniel, Bıyık, Erdem |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MILE: Model-based Intervention Learning
by: Korkmaz, Yigit, et al.
Published: (2025)
by: Korkmaz, Yigit, et al.
Published: (2025)
Actor-Free Continuous Control via Structurally Maximizable Q-Functions
by: Korkmaz, Yigit, et al.
Published: (2025)
by: Korkmaz, Yigit, et al.
Published: (2025)
Batch Active Learning of Reward Functions from Human Preferences
by: Bıyık, Erdem, et al.
Published: (2024)
by: Bıyık, Erdem, et al.
Published: (2024)
When a Robot is More Capable than a Human: Learning from Constrained Demonstrators
by: Li, Xinhu, et al.
Published: (2025)
by: Li, Xinhu, et al.
Published: (2025)
IMPACT: Intelligent Motion Planning with Acceptable Contact Trajectories via Vision-Language Models
by: Ling, Yiyang, et al.
Published: (2025)
by: Ling, Yiyang, et al.
Published: (2025)
A Generalized Acquisition Function for Preference-based Reward Learning
by: Ellis, Evan, et al.
Published: (2024)
by: Ellis, Evan, et al.
Published: (2024)
GABRIL: Gaze-Based Regularization for Mitigating Causal Confusion in Imitation Learning
by: Banayeeanzade, Amin, et al.
Published: (2025)
by: Banayeeanzade, Amin, et al.
Published: (2025)
Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
by: Liang, Anthony, et al.
Published: (2026)
by: Liang, Anthony, et al.
Published: (2026)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
by: Wang, Yufei, et al.
Published: (2024)
by: Wang, Yufei, et al.
Published: (2024)
CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations
by: Liang, Anthony, et al.
Published: (2025)
by: Liang, Anthony, et al.
Published: (2025)
Multi-Agent Inverse Q-Learning from Demonstrations
by: Haynam, Nathaniel, et al.
Published: (2025)
by: Haynam, Nathaniel, et al.
Published: (2025)
EXTRACT: Efficient Policy Learning by Extracting Transferable Robot Skills from Offline Data
by: Zhang, Jesse, et al.
Published: (2024)
by: Zhang, Jesse, et al.
Published: (2024)
Adaptive Querying for Reward Learning from Human Feedback
by: Anand, Yashwanthi, et al.
Published: (2024)
by: Anand, Yashwanthi, et al.
Published: (2024)
Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions
by: Jain, Ayush, et al.
Published: (2024)
by: Jain, Ayush, et al.
Published: (2024)
Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
by: Zhang, Rongtao, et al.
Published: (2026)
by: Zhang, Rongtao, et al.
Published: (2026)
Residual Reward Models for Preference-based Reinforcement Learning
by: Cao, Chenyang, et al.
Published: (2025)
by: Cao, Chenyang, et al.
Published: (2025)
RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences
by: Cheng, Jie, et al.
Published: (2024)
by: Cheng, Jie, et al.
Published: (2024)
PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies
by: Zhang, Jesse, et al.
Published: (2025)
by: Zhang, Jesse, et al.
Published: (2025)
Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented Imitation
by: Guo, Yihong, et al.
Published: (2024)
by: Guo, Yihong, et al.
Published: (2024)
D-CODA: Diffusion for Coordinated Dual-Arm Data Augmentation
by: Liu, I-Chun Arthur, et al.
Published: (2025)
by: Liu, I-Chun Arthur, et al.
Published: (2025)
Causal Flow Q-Learning for Robust Offline Reinforcement Learning
by: Li, Mingxuan, et al.
Published: (2026)
by: Li, Mingxuan, et al.
Published: (2026)
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
by: Poddar, Sriyash, et al.
Published: (2024)
by: Poddar, Sriyash, et al.
Published: (2024)
Developing Driving Strategies Efficiently: A Skill-Based Hierarchical Reinforcement Learning Approach
by: Gurses, Yigit, et al.
Published: (2023)
by: Gurses, Yigit, et al.
Published: (2023)
ROPA: Synthetic Robot Pose Generation for RGB-D Bimanual Data Augmentation
by: Chen, Jason, et al.
Published: (2025)
by: Chen, Jason, et al.
Published: (2025)
ViSaRL: Visual Reinforcement Learning Guided by Human Saliency
by: Liang, Anthony, et al.
Published: (2024)
by: Liang, Anthony, et al.
Published: (2024)
Reward Learning from Suboptimal Demonstrations with Applications in Surgical Electrocautery
by: Karimi, Zohre, et al.
Published: (2024)
by: Karimi, Zohre, et al.
Published: (2024)
Causal Action Influence Aware Counterfactual Data Augmentation
by: Urpí, Núria Armengol, et al.
Published: (2024)
by: Urpí, Núria Armengol, et al.
Published: (2024)
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
by: Ishihara, Yu, et al.
Published: (2025)
by: Ishihara, Yu, et al.
Published: (2025)
Confounding Robust Continuous Control via Automatic Reward Shaping
by: Juliani, Mateo, et al.
Published: (2026)
by: Juliani, Mateo, et al.
Published: (2026)
CaRL: Learning Scalable Planning Policies with Simple Rewards
by: Jaeger, Bernhard, et al.
Published: (2025)
by: Jaeger, Bernhard, et al.
Published: (2025)
TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
by: Liu, Yuyang, et al.
Published: (2025)
by: Liu, Yuyang, et al.
Published: (2025)
Object and Relation Centric Representations for Push Effect Prediction
by: Tekden, Ahmet E., et al.
Published: (2021)
by: Tekden, Ahmet E., et al.
Published: (2021)
Predictive Preference Learning from Human Interventions
by: Cai, Haoyuan, et al.
Published: (2025)
by: Cai, Haoyuan, et al.
Published: (2025)
Diffusion-Reward Adversarial Imitation Learning
by: Lai, Chun-Mao, et al.
Published: (2024)
by: Lai, Chun-Mao, et al.
Published: (2024)
Learning Causal Structure Distributions for Robust Planning
by: Murillo-Gonzalez, Alejandro, et al.
Published: (2025)
by: Murillo-Gonzalez, Alejandro, et al.
Published: (2025)
Reward-Punishment Reinforcement Learning with Maximum Entropy
by: Wang, Jiexin, et al.
Published: (2024)
by: Wang, Jiexin, et al.
Published: (2024)
Beyond Scalar Rewards: Distributional Reinforcement Learning with Preordered Objectives for Safe and Reliable Autonomous Driving
by: Abouelazm, Ahmed, et al.
Published: (2026)
by: Abouelazm, Ahmed, et al.
Published: (2026)
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
by: Singh, Anukriti, et al.
Published: (2025)
by: Singh, Anukriti, et al.
Published: (2025)
Subwords as Skills: Tokenization for Sparse-Reward Reinforcement Learning
by: Yunis, David, et al.
Published: (2023)
by: Yunis, David, et al.
Published: (2023)
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
by: Diaz-Bone, Leander, et al.
Published: (2025)
by: Diaz-Bone, Leander, et al.
Published: (2025)
Similar Items
-
MILE: Model-based Intervention Learning
by: Korkmaz, Yigit, et al.
Published: (2025) -
Actor-Free Continuous Control via Structurally Maximizable Q-Functions
by: Korkmaz, Yigit, et al.
Published: (2025) -
Batch Active Learning of Reward Functions from Human Preferences
by: Bıyık, Erdem, et al.
Published: (2024) -
When a Robot is More Capable than a Human: Learning from Constrained Demonstrators
by: Li, Xinhu, et al.
Published: (2025) -
IMPACT: Intelligent Motion Planning with Acceptable Contact Trajectories via Vision-Language Models
by: Ling, Yiyang, et al.
Published: (2025)