REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Chakraborty, Souradip, Singh, Anukriti, Bhaskar, Amisha, Tokekar, Pratap, Manocha, Dinesh, Bedi, Amrit Singh |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
by: Singh, Anukriti, et al.
Published: (2025)
by: Singh, Anukriti, et al.
Published: (2025)
Sketch-to-Skill: Bootstrapping Robot Learning with Human Drawn Trajectory Sketches
by: Yu, Peihong, et al.
Published: (2025)
by: Yu, Peihong, et al.
Published: (2025)
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
by: Barakat, Anas, et al.
Published: (2024)
by: Barakat, Anas, et al.
Published: (2024)
PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback
by: Chakraborty, Souradip, et al.
Published: (2023)
by: Chakraborty, Souradip, et al.
Published: (2023)
MaxMin-RLHF: Alignment with Diverse Human Preferences
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
Adaptive Visual Imitation Learning for Robotic Assisted Feeding Across Varied Bowl Configurations and Food Types
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
PRISM: Performer RS-IMLE for Single-pass Multisensory Imitation Learning
by: Bhaskar, Amisha, et al.
Published: (2026)
by: Bhaskar, Amisha, et al.
Published: (2026)
Learning Multi-Robot Coordination through Locality-Based Factorized Multi-Agent Actor-Critic Algorithm
by: Shek, Chak Lam, et al.
Published: (2025)
by: Shek, Chak Lam, et al.
Published: (2025)
Pre-Trained Masked Image Model for Mobile Robot Navigation
by: Sharma, Vishnu Dutt, et al.
Published: (2023)
by: Sharma, Vishnu Dutt, et al.
Published: (2023)
RL with Learnable Textual Feedback: A Bilevel Approach
by: Singh, Utsav, et al.
Published: (2026)
by: Singh, Utsav, et al.
Published: (2026)
AG-CVG: Coverage Planning with a Mobile Recharging UGV and an Energy-Constrained UAV
by: Karapetyan, Nare, et al.
Published: (2023)
by: Karapetyan, Nare, et al.
Published: (2023)
Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation
by: Patel, Bhrij, et al.
Published: (2023)
by: Patel, Bhrij, et al.
Published: (2023)
LANCAR: Leveraging Language for Context-Aware Robot Locomotion in Unstructured Environments
by: Shek, Chak Lam, et al.
Published: (2023)
by: Shek, Chak Lam, et al.
Published: (2023)
DMCA: Dense Multi-agent Navigation using Attention and Communication
by: Arul, Senthil Hariharan, et al.
Published: (2022)
by: Arul, Senthil Hariharan, et al.
Published: (2022)
Beyond Joint Demonstrations: Personalized Expert Guidance for Efficient Multi-Agent Reinforcement Learning
by: Yu, Peihong, et al.
Published: (2024)
by: Yu, Peihong, et al.
Published: (2024)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
Beyond Text: Utilizing Vocal Cues to Improve Decision Making in LLMs for Robot Navigation Tasks
by: Sun, Xingpeng, et al.
Published: (2024)
by: Sun, Xingpeng, et al.
Published: (2024)
On the Vulnerability of LLM/VLM-Controlled Robotics
by: Wu, Xiyang, et al.
Published: (2024)
by: Wu, Xiyang, et al.
Published: (2024)
IMRL: Integrating Visual, Physical, Temporal, and Geometric Representations for Enhanced Food Acquisition
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation
by: Singh, Anukriti, et al.
Published: (2025)
by: Singh, Anukriti, et al.
Published: (2025)
Personalized Embodied Navigation for Portable Object Finding
by: Dorbala, Vishnu Sashank, et al.
Published: (2024)
by: Dorbala, Vishnu Sashank, et al.
Published: (2024)
SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models
by: Wu, Xiyang, et al.
Published: (2026)
by: Wu, Xiyang, et al.
Published: (2026)
LAVA: Long-horizon Visual Action based Food Acquisition
by: Bhaskar, Amisha, et al.
Published: (2024)
by: Bhaskar, Amisha, et al.
Published: (2024)
Transfer Q Star: Principled Decoding for LLM Alignment
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
by: Barakat, Anas, et al.
Published: (2026)
by: Barakat, Anas, et al.
Published: (2026)
Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time Oracles
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
by: Ghosal, Soumya Suvra, et al.
Published: (2026)
by: Ghosal, Soumya Suvra, et al.
Published: (2026)
Improving Zero-Shot ObjectNav with Generative Communication
by: Dorbala, Vishnu Sashank, et al.
Published: (2024)
by: Dorbala, Vishnu Sashank, et al.
Published: (2024)
DIPPER: Direct Preference Optimization to Accelerate Primitive-Enabled Hierarchical Reinforcement Learning
by: Singh, Utsav, et al.
Published: (2024)
by: Singh, Utsav, et al.
Published: (2024)
Multi-LLM QA with Embodied Exploration
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
TrustNavGPT: Modeling Uncertainty to Improve Trustworthiness of Audio-Guided LLM-Based Robot Navigation
by: Sun, Xingpeng, et al.
Published: (2024)
by: Sun, Xingpeng, et al.
Published: (2024)
PLANRL: A Motion Planning and Imitation Learning Framework to Bootstrap Reinforcement Learning
by: Bhaskar, Amisha, et al.
Published: (2024)
by: Bhaskar, Amisha, et al.
Published: (2024)
MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
by: Zhai, Kevin, et al.
Published: (2025)
by: Zhai, Kevin, et al.
Published: (2025)
When to Localize? A Risk-Constrained Reinforcement Learning Approach
by: Shek, Chak Lam, et al.
Published: (2024)
by: Shek, Chak Lam, et al.
Published: (2024)
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
by: Trivedi, Prashant, et al.
Published: (2025)
by: Trivedi, Prashant, et al.
Published: (2025)
Formulating Reinforcement Learning for Human-Robot Collaboration through Off-Policy Evaluation
by: Singh, Saurav, et al.
Published: (2026)
by: Singh, Saurav, et al.
Published: (2026)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
When to Localize? A POMDP Approach
by: Williams, Troi, et al.
Published: (2024)
by: Williams, Troi, et al.
Published: (2024)
HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation
by: Seneviratne, Gershom, et al.
Published: (2025)
by: Seneviratne, Gershom, et al.
Published: (2025)
D2M2N: Decentralized Differentiable Memory-Enabled Mapping and Navigation for Multiple Robots
by: Ishat-E-Rabban, Md, et al.
Published: (2023)
by: Ishat-E-Rabban, Md, et al.
Published: (2023)
Similar Items
-
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
by: Singh, Anukriti, et al.
Published: (2025) -
Sketch-to-Skill: Bootstrapping Robot Learning with Human Drawn Trajectory Sketches
by: Yu, Peihong, et al.
Published: (2025) -
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
by: Barakat, Anas, et al.
Published: (2024) -
PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback
by: Chakraborty, Souradip, et al.
Published: (2023) -
MaxMin-RLHF: Alignment with Diverse Human Preferences
by: Chakraborty, Souradip, et al.
Published: (2024)