ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jiahui, Luo, Yusen, Anwar, Abrar, Sontakke, Sumedh Anand, Lim, Joseph J, Thomason, Jesse, Biyik, Erdem, Zhang, Jesse |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Contrast Sets for Evaluating Language-Guided Robot Policies
by: Anwar, Abrar, et al.
Published: (2024)
by: Anwar, Abrar, et al.
Published: (2024)
ViSaRL: Visual Reinforcement Learning Guided by Human Saliency
by: Liang, Anthony, et al.
Published: (2024)
by: Liang, Anthony, et al.
Published: (2024)
PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies
by: Zhang, Jesse, et al.
Published: (2025)
by: Zhang, Jesse, et al.
Published: (2025)
HAND Me the Data: Fast Robot Adaptation via Hand Path Retrieval
by: Hong, Matthew, et al.
Published: (2025)
by: Hong, Matthew, et al.
Published: (2025)
M3PT: A Transformer for Multimodal, Multi-Party Social Signal Prediction with Person-aware Blockwise Attention
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
Value Explicit Pretraining for Learning Transferable Representations
by: Lekkala, Kiran, et al.
Published: (2023)
by: Lekkala, Kiran, et al.
Published: (2023)
RobotFleet: An Open-Source Framework for Centralized Multi-Robot Task Planning
by: Gupta, Rohan, et al.
Published: (2025)
by: Gupta, Rohan, et al.
Published: (2025)
Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection
by: Anwar, Abrar, et al.
Published: (2025)
by: Anwar, Abrar, et al.
Published: (2025)
EXTRACT: Efficient Policy Learning by Extracting Transferable Robot Skills from Offline Data
by: Zhang, Jesse, et al.
Published: (2024)
by: Zhang, Jesse, et al.
Published: (2024)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
by: Mitra, Chancharik, et al.
Published: (2025)
by: Mitra, Chancharik, et al.
Published: (2025)
Generating Contextually-Relevant Navigation Instructions for Blind and Low Vision People
by: Merchant, Zain, et al.
Published: (2024)
by: Merchant, Zain, et al.
Published: (2024)
Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
by: Liang, Anthony, et al.
Published: (2026)
by: Liang, Anthony, et al.
Published: (2026)
SPRINT: Scalable Policy Pre-Training via Language Instruction Relabeling
by: Zhang, Jesse, et al.
Published: (2023)
by: Zhang, Jesse, et al.
Published: (2023)
Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding
by: Mitra, Chancharik, et al.
Published: (2023)
by: Mitra, Chancharik, et al.
Published: (2023)
Training robots with natural and lightweight human feedback
by: Erdem Bıyık
Published: (2026)
by: Erdem Bıyık
Published: (2026)
Batch Active Learning of Reward Functions from Human Preferences
by: Bıyık, Erdem, et al.
Published: (2024)
by: Bıyık, Erdem, et al.
Published: (2024)
Learning to Deliberate: Meta-policy Collaboration for Agentic LLMs with Multi-agent Reinforcement Learning
by: Yang, Wei, et al.
Published: (2025)
by: Yang, Wei, et al.
Published: (2025)
Adjust for Trust: Mitigating Trust-Induced Inappropriate Reliance on AI Assistance
by: Srinivasan, Tejas, et al.
Published: (2025)
by: Srinivasan, Tejas, et al.
Published: (2025)
When a Robot is More Capable than a Human: Learning from Constrained Demonstrators
by: Li, Xinhu, et al.
Published: (2025)
by: Li, Xinhu, et al.
Published: (2025)
CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations
by: Liang, Anthony, et al.
Published: (2025)
by: Liang, Anthony, et al.
Published: (2025)
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
by: Hwang, Minjune, et al.
Published: (2026)
by: Hwang, Minjune, et al.
Published: (2026)
MILE: Model-based Intervention Learning
by: Korkmaz, Yigit, et al.
Published: (2025)
by: Korkmaz, Yigit, et al.
Published: (2025)
Multi-Agent Inverse Q-Learning from Demonstrations
by: Haynam, Nathaniel, et al.
Published: (2025)
by: Haynam, Nathaniel, et al.
Published: (2025)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
by: Wang, Yufei, et al.
Published: (2024)
by: Wang, Yufei, et al.
Published: (2024)
Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions
by: Jain, Ayush, et al.
Published: (2024)
by: Jain, Ayush, et al.
Published: (2024)
Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM
by: Hu, Zizhao, et al.
Published: (2026)
by: Hu, Zizhao, et al.
Published: (2026)
TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP
by: Cai, Yuliang, et al.
Published: (2025)
by: Cai, Yuliang, et al.
Published: (2025)
Phonological Representation Learning for Isolated Signs Improves Out-of-Vocabulary Generalization
by: Kezar, Lee, et al.
Published: (2025)
by: Kezar, Lee, et al.
Published: (2025)
Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models
by: Hu, Zizhao, et al.
Published: (2025)
by: Hu, Zizhao, et al.
Published: (2025)
Iterative Formalization and Planning in Partially Observable Environments
by: Gong, Liancheng, et al.
Published: (2025)
by: Gong, Liancheng, et al.
Published: (2025)
Words that make SENSE: Sensorimotor Norms in Learned Lexical Token Representations
by: Gupta, Abhinav, et al.
Published: (2026)
by: Gupta, Abhinav, et al.
Published: (2026)
Do Localization Methods Actually Localize Memorized Data in LLMs? A Tale of Two Benchmarks
by: Chang, Ting-Yun, et al.
Published: (2023)
by: Chang, Ting-Yun, et al.
Published: (2023)
When Parts Are Greater Than Sums: Individual LLM Components Can Outperform Full Models
by: Chang, Ting-Yun, et al.
Published: (2024)
by: Chang, Ting-Yun, et al.
Published: (2024)
Why Do Some Inputs Break Low-Bit LLM Quantization?
by: Chang, Ting-Yun, et al.
Published: (2025)
by: Chang, Ting-Yun, et al.
Published: (2025)
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
by: He, Keyu, et al.
Published: (2025)
by: He, Keyu, et al.
Published: (2025)
A Generalized Acquisition Function for Preference-based Reward Learning
by: Ellis, Evan, et al.
Published: (2024)
by: Ellis, Evan, et al.
Published: (2024)
Judgelight: Trajectory-Level Post-Optimization for Multi-Agent Path Finding via Closed-Subwalk Collapsing
by: Tang, Yimin, et al.
Published: (2026)
by: Tang, Yimin, et al.
Published: (2026)
". . . And Gladly Teach."
by: Shera, Jesse H.
Published: (1978)
by: Shera, Jesse H.
Published: (1978)
Policy-Guided Search on Tree-of-Thoughts for Efficient Problem Solving with Bounded Language Model Queries
by: Pendurkar, Sumedh, et al.
Published: (2026)
by: Pendurkar, Sumedh, et al.
Published: (2026)
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
by: Pumacay, Wilbert, et al.
Published: (2024)
by: Pumacay, Wilbert, et al.
Published: (2024)
Similar Items
-
Contrast Sets for Evaluating Language-Guided Robot Policies
by: Anwar, Abrar, et al.
Published: (2024) -
ViSaRL: Visual Reinforcement Learning Guided by Human Saliency
by: Liang, Anthony, et al.
Published: (2024) -
PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies
by: Zhang, Jesse, et al.
Published: (2025) -
HAND Me the Data: Fast Robot Adaptation via Hand Path Retrieval
by: Hong, Matthew, et al.
Published: (2025) -
M3PT: A Transformer for Multimodal, Multi-Party Social Signal Prediction with Person-aware Blockwise Attention
by: Tang, Yiming, et al.
Published: (2025)