Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
Fuente:
arXiv
Salvato in:
| Autori principali: | Liang, Anthony, Korkmaz, Yigit, Zhang, Jiahui, Hwang, Minyoung, Anwar, Abrar, Kaushik, Sidhant, Shah, Aditya, Huang, Alex S., Zettlemoyer, Luke, Fox, Dieter, Xiang, Yu, Li, Anqi, Bobu, Andreea, Gupta, Abhishek, Tu, Stephen, Biyik, Erdem, Zhang, Jesse |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
di: Hwang, Minjune, et al.
Pubblicazione: (2026)
di: Hwang, Minjune, et al.
Pubblicazione: (2026)
MILE: Model-based Intervention Learning
di: Korkmaz, Yigit, et al.
Pubblicazione: (2025)
di: Korkmaz, Yigit, et al.
Pubblicazione: (2025)
Masked IRL: LLM-Guided Reward Disambiguation from Demonstrations and Language
di: Hwang, Minyoung, et al.
Pubblicazione: (2025)
di: Hwang, Minyoung, et al.
Pubblicazione: (2025)
ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations
di: Zhang, Jiahui, et al.
Pubblicazione: (2025)
di: Zhang, Jiahui, et al.
Pubblicazione: (2025)
Actor-Free Continuous Control via Structurally Maximizable Q-Functions
di: Korkmaz, Yigit, et al.
Pubblicazione: (2025)
di: Korkmaz, Yigit, et al.
Pubblicazione: (2025)
When a Robot is More Capable than a Human: Learning from Constrained Demonstrators
di: Li, Xinhu, et al.
Pubblicazione: (2025)
di: Li, Xinhu, et al.
Pubblicazione: (2025)
PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies
di: Zhang, Jesse, et al.
Pubblicazione: (2025)
di: Zhang, Jesse, et al.
Pubblicazione: (2025)
GIFT: Generalizing Intent for Flexible Test-Time Rewards
di: Amin, Fin, et al.
Pubblicazione: (2026)
di: Amin, Fin, et al.
Pubblicazione: (2026)
Robots That Know What to Ask: Recovering Misaligned Rewards through Targeted Explanations
di: Merker, Helena, et al.
Pubblicazione: (2026)
di: Merker, Helena, et al.
Pubblicazione: (2026)
Learning Contextually-Adaptive Rewards via Calibrated Features
di: Forsey-Smerek, Alexandra, et al.
Pubblicazione: (2025)
di: Forsey-Smerek, Alexandra, et al.
Pubblicazione: (2025)
Training robots with natural and lightweight human feedback
di: Erdem Bıyık
Pubblicazione: (2026)
di: Erdem Bıyık
Pubblicazione: (2026)
ViSaRL: Visual Reinforcement Learning Guided by Human Saliency
di: Liang, Anthony, et al.
Pubblicazione: (2024)
di: Liang, Anthony, et al.
Pubblicazione: (2024)
Batch Active Learning of Reward Functions from Human Preferences
di: Bıyık, Erdem, et al.
Pubblicazione: (2024)
di: Bıyık, Erdem, et al.
Pubblicazione: (2024)
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
di: Yasunaga, Michihiro, et al.
Pubblicazione: (2025)
di: Yasunaga, Michihiro, et al.
Pubblicazione: (2025)
HAND Me the Data: Fast Robot Adaptation via Hand Path Retrieval
di: Hong, Matthew, et al.
Pubblicazione: (2025)
di: Hong, Matthew, et al.
Pubblicazione: (2025)
M3PT: A Transformer for Multimodal, Multi-Party Social Signal Prediction with Person-aware Blockwise Attention
di: Tang, Yiming, et al.
Pubblicazione: (2025)
di: Tang, Yiming, et al.
Pubblicazione: (2025)
Contrast Sets for Evaluating Language-Guided Robot Policies
di: Anwar, Abrar, et al.
Pubblicazione: (2024)
di: Anwar, Abrar, et al.
Pubblicazione: (2024)
A Generalized Acquisition Function for Preference-based Reward Learning
di: Ellis, Evan, et al.
Pubblicazione: (2024)
di: Ellis, Evan, et al.
Pubblicazione: (2024)
Judgelight: Trajectory-Level Post-Optimization for Multi-Agent Path Finding via Closed-Subwalk Collapsing
di: Tang, Yimin, et al.
Pubblicazione: (2026)
di: Tang, Yimin, et al.
Pubblicazione: (2026)
Charged energy correlators in small systems with ALICE
di: Hwang, Minyoung
Pubblicazione: (2025)
di: Hwang, Minyoung
Pubblicazione: (2025)
Trajectory Improvement and Reward Learning from Comparative Language Feedback
di: Yang, Zhaojing, et al.
Pubblicazione: (2024)
di: Yang, Zhaojing, et al.
Pubblicazione: (2024)
Reprogramming Filamentous fd Viruses to Capture Copper Ions
di: Nuriye Korkmaz, et al.
Pubblicazione: (2024)
di: Nuriye Korkmaz, et al.
Pubblicazione: (2024)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
di: Wang, Yufei, et al.
Pubblicazione: (2024)
di: Wang, Yufei, et al.
Pubblicazione: (2024)
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
di: Ma, Rachel, et al.
Pubblicazione: (2025)
di: Ma, Rachel, et al.
Pubblicazione: (2025)
Goal Inference from Open-Ended Dialog
di: Ma, Rachel, et al.
Pubblicazione: (2024)
di: Ma, Rachel, et al.
Pubblicazione: (2024)
QuickLAP: Quick Language-Action Preference Learning for Semi-Autonomous Agents
di: Nader, Jordan Abi, et al.
Pubblicazione: (2025)
di: Nader, Jordan Abi, et al.
Pubblicazione: (2025)
Human-Guided Harm Recovery for Computer Use Agents
di: Li, Christy, et al.
Pubblicazione: (2026)
di: Li, Christy, et al.
Pubblicazione: (2026)
CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations
di: Liang, Anthony, et al.
Pubblicazione: (2025)
di: Liang, Anthony, et al.
Pubblicazione: (2025)
EXTRACT: Efficient Policy Learning by Extracting Transferable Robot Skills from Offline Data
di: Zhang, Jesse, et al.
Pubblicazione: (2024)
di: Zhang, Jesse, et al.
Pubblicazione: (2024)
GABRIL: Gaze-Based Regularization for Mitigating Causal Confusion in Imitation Learning
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2025)
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2025)
Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree
di: Jaggi, Harbani, et al.
Pubblicazione: (2024)
di: Jaggi, Harbani, et al.
Pubblicazione: (2024)
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
di: Damani, Mehul, et al.
Pubblicazione: (2024)
di: Damani, Mehul, et al.
Pubblicazione: (2024)
Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
di: Hu, Yushi, et al.
Pubblicazione: (2025)
di: Hu, Yushi, et al.
Pubblicazione: (2025)
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
di: Chen, Tong, et al.
Pubblicazione: (2025)
di: Chen, Tong, et al.
Pubblicazione: (2025)
Aligning Robot and Human Representations
di: Bobu, Andreea, et al.
Pubblicazione: (2023)
di: Bobu, Andreea, et al.
Pubblicazione: (2023)
Generating Contextually-Relevant Navigation Instructions for Blind and Low Vision People
di: Merchant, Zain, et al.
Pubblicazione: (2024)
di: Merchant, Zain, et al.
Pubblicazione: (2024)
RobotFleet: An Open-Source Framework for Centralized Multi-Robot Task Planning
di: Gupta, Rohan, et al.
Pubblicazione: (2025)
di: Gupta, Rohan, et al.
Pubblicazione: (2025)
Improving through Interaction: Searching Behavioral Representation Spaces with CMA-ES-IG
di: Dennler, Nathaniel, et al.
Pubblicazione: (2026)
di: Dennler, Nathaniel, et al.
Pubblicazione: (2026)
Coprocessor Actor Critic: A Model-Based Reinforcement Learning Approach For Adaptive Brain Stimulation
di: Pan, Michelle, et al.
Pubblicazione: (2024)
di: Pan, Michelle, et al.
Pubblicazione: (2024)
Multi-Agent Inverse Q-Learning from Demonstrations
di: Haynam, Nathaniel, et al.
Pubblicazione: (2025)
di: Haynam, Nathaniel, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
di: Hwang, Minjune, et al.
Pubblicazione: (2026) -
MILE: Model-based Intervention Learning
di: Korkmaz, Yigit, et al.
Pubblicazione: (2025) -
Masked IRL: LLM-Guided Reward Disambiguation from Demonstrations and Language
di: Hwang, Minyoung, et al.
Pubblicazione: (2025) -
ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations
di: Zhang, Jiahui, et al.
Pubblicazione: (2025) -
Actor-Free Continuous Control via Structurally Maximizable Q-Functions
di: Korkmaz, Yigit, et al.
Pubblicazione: (2025)