Policy Learning from Large Vision-Language Model Feedback without Reward Modeling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Luu, Tung M., Lee, Donghoon, Lee, Younghwan, Yoo, Chang D. |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models
par: Luu, Tung Minh, et autres
Publié: (2025)
par: Luu, Tung Minh, et autres
Publié: (2025)
Reward Generation via Large Vision-Language Model in Offline Reinforcement Learning
par: Lee, Younghwan, et autres
Publié: (2025)
par: Lee, Younghwan, et autres
Publié: (2025)
Sample Efficient Reinforcement Learning via Large Vision Language Model Distillation
par: Lee, Donghoon, et autres
Publié: (2025)
par: Lee, Donghoon, et autres
Publié: (2025)
Towards Robust Policy: Enhancing Offline Reinforcement Learning with Adversarial Attacks and Defenses
par: Nguyen, Thanh, et autres
Publié: (2024)
par: Nguyen, Thanh, et autres
Publié: (2024)
Predictive Coding for Decision Transformer
par: Luu, Tung M., et autres
Publié: (2024)
par: Luu, Tung M., et autres
Publié: (2024)
Uncertainty-Aware Rank-One MIMO Q Network Framework for Accelerated Offline Reinforcement Learning
par: Nguyen, Thanh, et autres
Publié: (2026)
par: Nguyen, Thanh, et autres
Publié: (2026)
ELEMENTAL: Interactive Learning from Demonstrations and Vision-Language Models for Reward Design in Robotics
par: Chen, Letian, et autres
Publié: (2024)
par: Chen, Letian, et autres
Publié: (2024)
DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
par: Lee, Vint, et autres
Publié: (2023)
par: Lee, Vint, et autres
Publié: (2023)
Rewarding DINO: Predicting Dense Rewards with Vision Foundation Models
par: Krack, Pierre, et autres
Publié: (2026)
par: Krack, Pierre, et autres
Publié: (2026)
Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
par: Zheng, Qinqing, et autres
Publié: (2024)
par: Zheng, Qinqing, et autres
Publié: (2024)
Real-World Offline Reinforcement Learning from Vision Language Model Feedback
par: Venkataraman, Sreyas, et autres
Publié: (2024)
par: Venkataraman, Sreyas, et autres
Publié: (2024)
LAPP: Large Language Model Feedback for Preference-Driven Reinforcement Learning
par: Jian, Pingcheng, et autres
Publié: (2025)
par: Jian, Pingcheng, et autres
Publié: (2025)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
par: Wang, Yufei, et autres
Publié: (2024)
par: Wang, Yufei, et autres
Publié: (2024)
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution
par: Huang, Changxin, et autres
Publié: (2024)
par: Huang, Changxin, et autres
Publié: (2024)
Not Only Rewards But Also Constraints: Applications on Legged Robot Locomotion
par: Kim, Yunho, et autres
Publié: (2023)
par: Kim, Yunho, et autres
Publié: (2023)
VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback
par: Bi, Jianxin, et autres
Publié: (2025)
par: Bi, Jianxin, et autres
Publié: (2025)
Learning a High-quality Robotic Wiping Policy Using Systematic Reward Analysis and Visual-Language Model Based Curriculum
par: Liu, Yihong, et autres
Publié: (2025)
par: Liu, Yihong, et autres
Publié: (2025)
Uncertainty-Aware Deployment of Pre-trained Language-Conditioned Imitation Learning Policies
par: Wu, Bo, et autres
Publié: (2024)
par: Wu, Bo, et autres
Publié: (2024)
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
par: Ishihara, Yu, et autres
Publié: (2025)
par: Ishihara, Yu, et autres
Publié: (2025)
Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
par: Lyu, Mingyang, et autres
Publié: (2025)
par: Lyu, Mingyang, et autres
Publié: (2025)
Adaptive Querying for Reward Learning from Human Feedback
par: Anand, Yashwanthi, et autres
Publié: (2024)
par: Anand, Yashwanthi, et autres
Publié: (2024)
REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
par: Chakraborty, Souradip, et autres
Publié: (2023)
par: Chakraborty, Souradip, et autres
Publié: (2023)
Few-Shot Vision-Language Action-Incremental Policy Learning
par: Song, Mingchen, et autres
Publié: (2025)
par: Song, Mingchen, et autres
Publié: (2025)
Learning Generalizable Visuomotor Policy through Dynamics-Alignment
par: Lee, Dohyeok, et autres
Publié: (2025)
par: Lee, Dohyeok, et autres
Publié: (2025)
Learning Emergent Gaits with Decentralized Phase Oscillators: on the role of Observations, Rewards, and Feedback
par: Zhang, Jenny, et autres
Publié: (2024)
par: Zhang, Jenny, et autres
Publié: (2024)
Habilis-$β$: A Fast-Motion and Long-Lasting On-Device Vision-Language-Action Model
par: Robotics, Tommoro, et autres
Publié: (2026)
par: Robotics, Tommoro, et autres
Publié: (2026)
TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
par: Im, Hokyun, et autres
Publié: (2025)
par: Im, Hokyun, et autres
Publié: (2025)
Enhancing Robotic Manipulation with AI Feedback from Multimodal Large Language Models
par: Liu, Jinyi, et autres
Publié: (2024)
par: Liu, Jinyi, et autres
Publié: (2024)
LearningFlow: Automated Policy Learning Workflow for Urban Driving with Large Language Models
par: Peng, Zengqi, et autres
Publié: (2025)
par: Peng, Zengqi, et autres
Publié: (2025)
APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model
par: Lu, Yuanjie, et autres
Publié: (2026)
par: Lu, Yuanjie, et autres
Publié: (2026)
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
par: Hwang, Minjune, et autres
Publié: (2026)
par: Hwang, Minjune, et autres
Publié: (2026)
RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control
par: Yue, Junpeng, et autres
Publié: (2025)
par: Yue, Junpeng, et autres
Publié: (2025)
Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models
par: Zhang, Zhilong, et autres
Publié: (2026)
par: Zhang, Zhilong, et autres
Publié: (2026)
Self-Supervised Curriculum Generation for Autonomous Reinforcement Learning without Task-Specific Knowledge
par: Lee, Sang-Hyun, et autres
Publié: (2023)
par: Lee, Sang-Hyun, et autres
Publié: (2023)
Eureka: Human-Level Reward Design via Coding Large Language Models
par: Ma, Yecheng Jason, et autres
Publié: (2023)
par: Ma, Yecheng Jason, et autres
Publié: (2023)
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
par: Singh, Anukriti, et autres
Publié: (2025)
par: Singh, Anukriti, et autres
Publié: (2025)
Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
par: Xie, Tianbao, et autres
Publié: (2023)
par: Xie, Tianbao, et autres
Publié: (2023)
The Lie We Tell: Correcting the Euclidean Fallacy in Vision Language Action Policies via Score Matching on Tangent Space
par: Chuang, Bing-Cheng, et autres
Publié: (2026)
par: Chuang, Bing-Cheng, et autres
Publié: (2026)
Chain of Uncertain Rewards with Large Language Models for Reinforcement Learning
par: Mo, Shentong
Publié: (2026)
par: Mo, Shentong
Publié: (2026)
STRIDE: Automating Reward Design, Deep Reinforcement Learning Training and Feedback Optimization in Humanoid Robotics Locomotion
par: Wu, Zhenwei, et autres
Publié: (2025)
par: Wu, Zhenwei, et autres
Publié: (2025)
Documents similaires
-
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models
par: Luu, Tung Minh, et autres
Publié: (2025) -
Reward Generation via Large Vision-Language Model in Offline Reinforcement Learning
par: Lee, Younghwan, et autres
Publié: (2025) -
Sample Efficient Reinforcement Learning via Large Vision Language Model Distillation
par: Lee, Donghoon, et autres
Publié: (2025) -
Towards Robust Policy: Enhancing Offline Reinforcement Learning with Adversarial Attacks and Defenses
par: Nguyen, Thanh, et autres
Publié: (2024) -
Predictive Coding for Decision Transformer
par: Luu, Tung M., et autres
Publié: (2024)