Trajectory Improvement and Reward Learning from Comparative Language Feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Zhaojing, Jun, Miru, Tien, Jeremy, Russell, Stuart J., Dragan, Anca, Bıyık, Erdem |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Generalized Acquisition Function for Preference-based Reward Learning
von: Ellis, Evan, et al.
Veröffentlicht: (2024)
von: Ellis, Evan, et al.
Veröffentlicht: (2024)
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
von: Hwang, Minjune, et al.
Veröffentlicht: (2026)
von: Hwang, Minjune, et al.
Veröffentlicht: (2026)
When a Robot is More Capable than a Human: Learning from Constrained Demonstrators
von: Li, Xinhu, et al.
Veröffentlicht: (2025)
von: Li, Xinhu, et al.
Veröffentlicht: (2025)
Batch Active Learning of Reward Functions from Human Preferences
von: Bıyık, Erdem, et al.
Veröffentlicht: (2024)
von: Bıyık, Erdem, et al.
Veröffentlicht: (2024)
Judgelight: Trajectory-Level Post-Optimization for Multi-Agent Path Finding via Closed-Subwalk Collapsing
von: Tang, Yimin, et al.
Veröffentlicht: (2026)
von: Tang, Yimin, et al.
Veröffentlicht: (2026)
MILE: Model-based Intervention Learning
von: Korkmaz, Yigit, et al.
Veröffentlicht: (2025)
von: Korkmaz, Yigit, et al.
Veröffentlicht: (2025)
ViSaRL: Visual Reinforcement Learning Guided by Human Saliency
von: Liang, Anthony, et al.
Veröffentlicht: (2024)
von: Liang, Anthony, et al.
Veröffentlicht: (2024)
IMPACT: Intelligent Motion Planning with Acceptable Contact Trajectories via Vision-Language Models
von: Ling, Yiyang, et al.
Veröffentlicht: (2025)
von: Ling, Yiyang, et al.
Veröffentlicht: (2025)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
von: Wang, Yufei, et al.
Veröffentlicht: (2024)
von: Wang, Yufei, et al.
Veröffentlicht: (2024)
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2024)
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2024)
ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations
von: Zhang, Jiahui, et al.
Veröffentlicht: (2025)
von: Zhang, Jiahui, et al.
Veröffentlicht: (2025)
SyncTwin: Fast Digital Twin Construction and Synchronization for Safe Robotic Manipulation
von: Huang, Ruopeng, et al.
Veröffentlicht: (2026)
von: Huang, Ruopeng, et al.
Veröffentlicht: (2026)
GABRIL: Gaze-Based Regularization for Mitigating Causal Confusion in Imitation Learning
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2025)
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2025)
Coprocessor Actor Critic: A Model-Based Reinforcement Learning Approach For Adaptive Brain Stimulation
von: Pan, Michelle, et al.
Veröffentlicht: (2024)
von: Pan, Michelle, et al.
Veröffentlicht: (2024)
Value Explicit Pretraining for Learning Transferable Representations
von: Lekkala, Kiran, et al.
Veröffentlicht: (2023)
von: Lekkala, Kiran, et al.
Veröffentlicht: (2023)
Multi-Agent Inverse Q-Learning from Demonstrations
von: Haynam, Nathaniel, et al.
Veröffentlicht: (2025)
von: Haynam, Nathaniel, et al.
Veröffentlicht: (2025)
Actor-Free Continuous Control via Structurally Maximizable Q-Functions
von: Korkmaz, Yigit, et al.
Veröffentlicht: (2025)
von: Korkmaz, Yigit, et al.
Veröffentlicht: (2025)
CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations
von: Liang, Anthony, et al.
Veröffentlicht: (2025)
von: Liang, Anthony, et al.
Veröffentlicht: (2025)
Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
von: Liang, Anthony, et al.
Veröffentlicht: (2026)
von: Liang, Anthony, et al.
Veröffentlicht: (2026)
EXTRACT: Efficient Policy Learning by Extracting Transferable Robot Skills from Offline Data
von: Zhang, Jesse, et al.
Veröffentlicht: (2024)
von: Zhang, Jesse, et al.
Veröffentlicht: (2024)
AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations
von: Gong, Litian, et al.
Veröffentlicht: (2025)
von: Gong, Litian, et al.
Veröffentlicht: (2025)
RAILGUN: A Unified Convolutional Policy for Multi-Agent Path Finding Across Different Environments and Tasks
von: Tang, Yimin, et al.
Veröffentlicht: (2025)
von: Tang, Yimin, et al.
Veröffentlicht: (2025)
HAND Me the Data: Fast Robot Adaptation via Hand Path Retrieval
von: Hong, Matthew, et al.
Veröffentlicht: (2025)
von: Hong, Matthew, et al.
Veröffentlicht: (2025)
AI Alignment with Changing and Influenceable Reward Functions
von: Carroll, Micah, et al.
Veröffentlicht: (2024)
von: Carroll, Micah, et al.
Veröffentlicht: (2024)
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling
von: Luu, Tung M., et al.
Veröffentlicht: (2025)
von: Luu, Tung M., et al.
Veröffentlicht: (2025)
Bridging RL Theory and Practice with the Effective Horizon
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2023)
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2023)
Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions
von: Jain, Ayush, et al.
Veröffentlicht: (2024)
von: Jain, Ayush, et al.
Veröffentlicht: (2024)
Adaptive Querying for Reward Learning from Human Feedback
von: Anand, Yashwanthi, et al.
Veröffentlicht: (2024)
von: Anand, Yashwanthi, et al.
Veröffentlicht: (2024)
Temporal Representation Alignment: Successor Features Enable Emergent Compositionality in Robot Instruction Following
von: Myers, Vivek, et al.
Veröffentlicht: (2025)
von: Myers, Vivek, et al.
Veröffentlicht: (2025)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
von: Lang, Leon, et al.
Veröffentlicht: (2024)
von: Lang, Leon, et al.
Veröffentlicht: (2024)
Training robots with natural and lightweight human feedback
von: Erdem Bıyık
Veröffentlicht: (2026)
von: Erdem Bıyık
Veröffentlicht: (2026)
Trajectory Planning for Autonomous Vehicle Using Iterative Reward Prediction in Reinforcement Learning
von: Park, Hyunwoo
Veröffentlicht: (2024)
von: Park, Hyunwoo
Veröffentlicht: (2024)
Toward Grounded Commonsense Reasoning
von: Kwon, Minae, et al.
Veröffentlicht: (2023)
von: Kwon, Minae, et al.
Veröffentlicht: (2023)
PRIMT: Preference-based Reinforcement Learning with Multimodal Feedback and Trajectory Synthesis from Foundation Models
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
Aligning Robot and Human Representations
von: Bobu, Andreea, et al.
Veröffentlicht: (2023)
von: Bobu, Andreea, et al.
Veröffentlicht: (2023)
REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2023)
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2023)
Collision Avoidance and Navigation for a Quadrotor Swarm Using End-to-end Deep Reinforcement Learning
von: Huang, Zhehui, et al.
Veröffentlicht: (2023)
von: Huang, Zhehui, et al.
Veröffentlicht: (2023)
Rank2Reward: Learning Shaped Reward Functions from Passive Video
von: Yang, Daniel, et al.
Veröffentlicht: (2024)
von: Yang, Daniel, et al.
Veröffentlicht: (2024)
Cooperative Inverse Reinforcement Learning
von: Hadfield-Menell, Dylan, et al.
Veröffentlicht: (2016)
von: Hadfield-Menell, Dylan, et al.
Veröffentlicht: (2016)
Learning Emergent Gaits with Decentralized Phase Oscillators: on the role of Observations, Rewards, and Feedback
von: Zhang, Jenny, et al.
Veröffentlicht: (2024)
von: Zhang, Jenny, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Generalized Acquisition Function for Preference-based Reward Learning
von: Ellis, Evan, et al.
Veröffentlicht: (2024) -
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
von: Hwang, Minjune, et al.
Veröffentlicht: (2026) -
When a Robot is More Capable than a Human: Learning from Constrained Demonstrators
von: Li, Xinhu, et al.
Veröffentlicht: (2025) -
Batch Active Learning of Reward Functions from Human Preferences
von: Bıyık, Erdem, et al.
Veröffentlicht: (2024) -
Judgelight: Trajectory-Level Post-Optimization for Multi-Agent Path Finding via Closed-Subwalk Collapsing
von: Tang, Yimin, et al.
Veröffentlicht: (2026)