Residual Reward Models for Preference-based Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Cao, Chenyang, Rogel-García, Miguel, Nabail, Mohamed, Wang, Xueqian, Rhinehart, Nicholas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DR-MPC: Deep Residual Model Predictive Control for Real-world Social Navigation
por: Han, James R., et al.
Publicado: (2024)
por: Han, James R., et al.
Publicado: (2024)
Offline Goal-Conditioned Reinforcement Learning for Safety-Critical Tasks with Recovery Policy
por: Cao, Chenyang, et al.
Publicado: (2024)
por: Cao, Chenyang, et al.
Publicado: (2024)
RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences
por: Cheng, Jie, et al.
Publicado: (2024)
por: Cheng, Jie, et al.
Publicado: (2024)
A Generalized Acquisition Function for Preference-based Reward Learning
por: Ellis, Evan, et al.
Publicado: (2024)
por: Ellis, Evan, et al.
Publicado: (2024)
AutoWorld: Scaling Multi-Agent Traffic Simulation with Self-Supervised World Models
por: Pourkeshavatz, Mozhgan, et al.
Publicado: (2026)
por: Pourkeshavatz, Mozhgan, et al.
Publicado: (2026)
Rating-based Reinforcement Learning
por: White, Devin, et al.
Publicado: (2023)
por: White, Devin, et al.
Publicado: (2023)
DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
por: Lee, Vint, et al.
Publicado: (2023)
por: Lee, Vint, et al.
Publicado: (2023)
Reward-Punishment Reinforcement Learning with Maximum Entropy
por: Wang, Jiexin, et al.
Publicado: (2024)
por: Wang, Jiexin, et al.
Publicado: (2024)
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
por: Ishihara, Yu, et al.
Publicado: (2025)
por: Ishihara, Yu, et al.
Publicado: (2025)
Batch Active Learning of Reward Functions from Human Preferences
por: Bıyık, Erdem, et al.
Publicado: (2024)
por: Bıyık, Erdem, et al.
Publicado: (2024)
Accelerating Residual Reinforcement Learning with Uncertainty Estimation
por: Dodeja, Lakshita, et al.
Publicado: (2025)
por: Dodeja, Lakshita, et al.
Publicado: (2025)
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
por: Hwang, Minjune, et al.
Publicado: (2026)
por: Hwang, Minjune, et al.
Publicado: (2026)
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
por: Diaz-Bone, Leander, et al.
Publicado: (2025)
por: Diaz-Bone, Leander, et al.
Publicado: (2025)
Subwords as Skills: Tokenization for Sparse-Reward Reinforcement Learning
por: Yunis, David, et al.
Publicado: (2023)
por: Yunis, David, et al.
Publicado: (2023)
Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
por: Xie, Tianbao, et al.
Publicado: (2023)
por: Xie, Tianbao, et al.
Publicado: (2023)
Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented Imitation
por: Guo, Yihong, et al.
Publicado: (2024)
por: Guo, Yihong, et al.
Publicado: (2024)
A Review of Reward Functions for Reinforcement Learning in the context of Autonomous Driving
por: Abouelazm, Ahmed, et al.
Publicado: (2024)
por: Abouelazm, Ahmed, et al.
Publicado: (2024)
Efficient Preference-Based Reinforcement Learning: Randomized Exploration Meets Experimental Design
por: Schlaginhaufen, Andreas, et al.
Publicado: (2025)
por: Schlaginhaufen, Andreas, et al.
Publicado: (2025)
Exploiting Symmetry in Dynamics for Model-Based Reinforcement Learning with Asymmetric Rewards
por: Sonmez, Yasin, et al.
Publicado: (2024)
por: Sonmez, Yasin, et al.
Publicado: (2024)
Beyond Scalar Rewards: Distributional Reinforcement Learning with Preordered Objectives for Safe and Reliable Autonomous Driving
por: Abouelazm, Ahmed, et al.
Publicado: (2026)
por: Abouelazm, Ahmed, et al.
Publicado: (2026)
Diffusion-Reward Adversarial Imitation Learning
por: Lai, Chun-Mao, et al.
Publicado: (2024)
por: Lai, Chun-Mao, et al.
Publicado: (2024)
Research on Autonomous Robots Navigation based on Reinforcement Learning
por: Wang, Zixiang, et al.
Publicado: (2024)
por: Wang, Zixiang, et al.
Publicado: (2024)
Innate-Values-driven Reinforcement Learning based Cognitive Modeling
por: Yang, Qin
Publicado: (2024)
por: Yang, Qin
Publicado: (2024)
Tactical Decision Making for Autonomous Trucks by Deep Reinforcement Learning with Total Cost of Operation Based Reward
por: Pathare, Deepthi, et al.
Publicado: (2024)
por: Pathare, Deepthi, et al.
Publicado: (2024)
From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models
por: Gumbsch, Christian, et al.
Publicado: (2026)
por: Gumbsch, Christian, et al.
Publicado: (2026)
LORD: Large Models based Opposite Reward Design for Autonomous Driving
por: Ye, Xin, et al.
Publicado: (2024)
por: Ye, Xin, et al.
Publicado: (2024)
PPNet: A Two-Stage Neural Network for End-to-end Path Planning
por: Meng, Qinglong, et al.
Publicado: (2024)
por: Meng, Qinglong, et al.
Publicado: (2024)
Temporal Distance-aware Transition Augmentation for Offline Model-based Reinforcement Learning
por: Lee, Dongsu, et al.
Publicado: (2025)
por: Lee, Dongsu, et al.
Publicado: (2025)
RDAR: Reward-Driven Agent Relevance Estimation for Autonomous Driving
por: Bosio, Carlo, et al.
Publicado: (2025)
por: Bosio, Carlo, et al.
Publicado: (2025)
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
por: Poddar, Sriyash, et al.
Publicado: (2024)
por: Poddar, Sriyash, et al.
Publicado: (2024)
Sample-Efficient Preference-based Reinforcement Learning with Dynamics Aware Rewards
por: Metcalf, Katherine, et al.
Publicado: (2024)
por: Metcalf, Katherine, et al.
Publicado: (2024)
WOMBET: World Model-based Experience Transfer for Robust and Sample-efficient Reinforcement Learning
por: Kim, Mintae, et al.
Publicado: (2026)
por: Kim, Mintae, et al.
Publicado: (2026)
Learning to Recover: Dynamic Reward Shaping with Wheel-Leg Coordination for Fallen Robots
por: Deng, Boyuan, et al.
Publicado: (2025)
por: Deng, Boyuan, et al.
Publicado: (2025)
Model-Based Reinforcement Learning with Multi-Task Offline Pretraining
por: Pan, Minting, et al.
Publicado: (2023)
por: Pan, Minting, et al.
Publicado: (2023)
Redistributing Rewards Across Time and Agents for Multi-Agent Reinforcement Learning
por: Kapoor, Aditya, et al.
Publicado: (2025)
por: Kapoor, Aditya, et al.
Publicado: (2025)
TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
por: Liu, Yuyang, et al.
Publicado: (2025)
por: Liu, Yuyang, et al.
Publicado: (2025)
Scaling Is All You Need: Autonomous Driving with JAX-Accelerated Reinforcement Learning
por: Harmel, Moritz, et al.
Publicado: (2023)
por: Harmel, Moritz, et al.
Publicado: (2023)
Physics-model-guided Worst-case Sampling for Safe Reinforcement Learning
por: Cao, Hongpeng, et al.
Publicado: (2024)
por: Cao, Hongpeng, et al.
Publicado: (2024)
Chain of Uncertain Rewards with Large Language Models for Reinforcement Learning
por: Mo, Shentong
Publicado: (2026)
por: Mo, Shentong
Publicado: (2026)
Adaptive Querying for Reward Learning from Human Feedback
por: Anand, Yashwanthi, et al.
Publicado: (2024)
por: Anand, Yashwanthi, et al.
Publicado: (2024)
Ejemplares similares
-
DR-MPC: Deep Residual Model Predictive Control for Real-world Social Navigation
por: Han, James R., et al.
Publicado: (2024) -
Offline Goal-Conditioned Reinforcement Learning for Safety-Critical Tasks with Recovery Policy
por: Cao, Chenyang, et al.
Publicado: (2024) -
RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences
por: Cheng, Jie, et al.
Publicado: (2024) -
A Generalized Acquisition Function for Preference-based Reward Learning
por: Ellis, Evan, et al.
Publicado: (2024) -
AutoWorld: Scaling Multi-Agent Traffic Simulation with Self-Supervised World Models
por: Pourkeshavatz, Mozhgan, et al.
Publicado: (2026)