The Dark Side of Rich Rewards: Understanding and Mitigating Noise in VLM Rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Sukai, Liu, Shu-Wei, Lipovetzky, Nir, Cohn, Trevor |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Planning in the Dark: LLM-Symbolic Planning Pipeline without Experts
by: Huang, Sukai, et al.
Published: (2024)
by: Huang, Sukai, et al.
Published: (2024)
Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan Generation
by: Huang, Sukai, et al.
Published: (2024)
by: Huang, Sukai, et al.
Published: (2024)
From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models
by: Gumbsch, Christian, et al.
Published: (2026)
by: Gumbsch, Christian, et al.
Published: (2026)
ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation
by: Feng, Youhe, et al.
Published: (2026)
by: Feng, Youhe, et al.
Published: (2026)
Rewarding DINO: Predicting Dense Rewards with Vision Foundation Models
by: Krack, Pierre, et al.
Published: (2026)
by: Krack, Pierre, et al.
Published: (2026)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
by: Zhang, Chen Bo Calvin, et al.
Published: (2024)
by: Zhang, Chen Bo Calvin, et al.
Published: (2024)
Diffusion Reward: Learning Rewards via Conditional Video Diffusion
by: Huang, Tao, et al.
Published: (2023)
by: Huang, Tao, et al.
Published: (2023)
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
by: Ishihara, Yu, et al.
Published: (2025)
by: Ishihara, Yu, et al.
Published: (2025)
Reward Machine Inference for Robotic Manipulation
by: Baert, Mattijs, et al.
Published: (2024)
by: Baert, Mattijs, et al.
Published: (2024)
Uncertainty-aware Reward Design Process
by: Yang, Yang, et al.
Published: (2025)
by: Yang, Yang, et al.
Published: (2025)
Curriculum Reinforcement Learning for Complex Reward Functions
by: Freitag, Kilian, et al.
Published: (2024)
by: Freitag, Kilian, et al.
Published: (2024)
On-Robot Reinforcement Learning with Goal-Contrastive Rewards
by: Biza, Ondrej, et al.
Published: (2024)
by: Biza, Ondrej, et al.
Published: (2024)
TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
by: Liu, Yuyang, et al.
Published: (2025)
by: Liu, Yuyang, et al.
Published: (2025)
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution
by: Huang, Changxin, et al.
Published: (2024)
by: Huang, Changxin, et al.
Published: (2024)
Video2Reward: Generating Reward Function from Videos for Legged Robot Behavior Learning
by: Zeng, Runhao, et al.
Published: (2024)
by: Zeng, Runhao, et al.
Published: (2024)
Revisiting Sparse Rewards for Goal-Reaching Reinforcement Learning
by: Vasan, Gautham, et al.
Published: (2024)
by: Vasan, Gautham, et al.
Published: (2024)
Reward Redistribution via Gaussian Process Likelihood Estimation
by: Xiao, Minheng, et al.
Published: (2025)
by: Xiao, Minheng, et al.
Published: (2025)
Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
by: Xie, Tianbao, et al.
Published: (2023)
by: Xie, Tianbao, et al.
Published: (2023)
Robot Policy Learning with Temporal Optimal Transport Reward
by: Fu, Yuwei, et al.
Published: (2024)
by: Fu, Yuwei, et al.
Published: (2024)
Generalization in Deep Reinforcement Learning for Robotic Navigation by Reward Shaping
by: Miranda, Victor R. F., et al.
Published: (2022)
by: Miranda, Victor R. F., et al.
Published: (2022)
Dual-Granularity Contrastive Reward via Generated Episodic Guidance for Efficient Embodied RL
by: Liu, Xin, et al.
Published: (2026)
by: Liu, Xin, et al.
Published: (2026)
Diffusion-Reward Adversarial Imitation Learning
by: Lai, Chun-Mao, et al.
Published: (2024)
by: Lai, Chun-Mao, et al.
Published: (2024)
Adaptive Teaching in Heterogeneous Agents: Balancing Surprise in Sparse Reward Scenarios
by: Clark, Emma, et al.
Published: (2024)
by: Clark, Emma, et al.
Published: (2024)
TopoNav: Topological Navigation for Efficient Exploration in Sparse Reward Environments
by: Hossain, Jumman, et al.
Published: (2024)
by: Hossain, Jumman, et al.
Published: (2024)
Bridging the Human to Robot Dexterity Gap through Object-Oriented Rewards
by: Guzey, Irmak, et al.
Published: (2024)
by: Guzey, Irmak, et al.
Published: (2024)
Reward Prediction Error Prioritisation in Experience Replay: The RPE-PER Method
by: Yamani, Hoda, et al.
Published: (2025)
by: Yamani, Hoda, et al.
Published: (2025)
ReLAM: Learning Anticipation Model for Rewarding Visual Robotic Manipulation
by: Tang, Nan, et al.
Published: (2025)
by: Tang, Nan, et al.
Published: (2025)
Momentum Based Reward Design for Low Emission Traffic Signal Control
by: Mundane, Chinmay, et al.
Published: (2026)
by: Mundane, Chinmay, et al.
Published: (2026)
DrS: Learning Reusable Dense Rewards for Multi-Stage Tasks
by: Mu, Tongzhou, et al.
Published: (2024)
by: Mu, Tongzhou, et al.
Published: (2024)
Beware Untrusted Simulators -- Reward-Free Backdoor Attacks in Reinforcement Learning
by: Rathbun, Ethan, et al.
Published: (2026)
by: Rathbun, Ethan, et al.
Published: (2026)
Learning Emergent Gaits with Decentralized Phase Oscillators: on the role of Observations, Rewards, and Feedback
by: Zhang, Jenny, et al.
Published: (2024)
by: Zhang, Jenny, et al.
Published: (2024)
Average-Reward Maximum Entropy Reinforcement Learning for Underactuated Double Pendulum Tasks
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
STRIDE: Automating Reward Design, Deep Reinforcement Learning Training and Feedback Optimization in Humanoid Robotics Locomotion
by: Wu, Zhenwei, et al.
Published: (2025)
by: Wu, Zhenwei, et al.
Published: (2025)
Reward-Punishment Reinforcement Learning with Maximum Entropy
by: Wang, Jiexin, et al.
Published: (2024)
by: Wang, Jiexin, et al.
Published: (2024)
ELEMENTAL: Interactive Learning from Demonstrations and Vision-Language Models for Reward Design in Robotics
by: Chen, Letian, et al.
Published: (2024)
by: Chen, Letian, et al.
Published: (2024)
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning
by: Venugopal, Aravind, et al.
Published: (2026)
by: Venugopal, Aravind, et al.
Published: (2026)
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling
by: Luu, Tung M., et al.
Published: (2025)
by: Luu, Tung M., et al.
Published: (2025)
REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
by: Chakraborty, Souradip, et al.
Published: (2023)
by: Chakraborty, Souradip, et al.
Published: (2023)
Decoupling Task and Behavior: A Two-Stage Reward Curriculum in Reinforcement Learning for Robotics
by: Freitag, Kilian, et al.
Published: (2026)
by: Freitag, Kilian, et al.
Published: (2026)
LORD: Large Models based Opposite Reward Design for Autonomous Driving
by: Ye, Xin, et al.
Published: (2024)
by: Ye, Xin, et al.
Published: (2024)
Similar Items
-
Planning in the Dark: LLM-Symbolic Planning Pipeline without Experts
by: Huang, Sukai, et al.
Published: (2024) -
Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan Generation
by: Huang, Sukai, et al.
Published: (2024) -
From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models
by: Gumbsch, Christian, et al.
Published: (2026) -
ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation
by: Feng, Youhe, et al.
Published: (2026) -
Rewarding DINO: Predicting Dense Rewards with Vision Foundation Models
by: Krack, Pierre, et al.
Published: (2026)