Masked IRL: LLM-Guided Reward Disambiguation from Demonstrations and Language
Fuente:
arXiv
Salvato in:
| Autori principali: | Hwang, Minyoung, Forsey-Smerek, Alexandra, Dennler, Nathaniel, Bobu, Andreea |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning Contextually-Adaptive Rewards via Calibrated Features
di: Forsey-Smerek, Alexandra, et al.
Pubblicazione: (2025)
di: Forsey-Smerek, Alexandra, et al.
Pubblicazione: (2025)
QuickLAP: Quick Language-Action Preference Learning for Semi-Autonomous Agents
di: Nader, Jordan Abi, et al.
Pubblicazione: (2025)
di: Nader, Jordan Abi, et al.
Pubblicazione: (2025)
GIFT: Generalizing Intent for Flexible Test-Time Rewards
di: Amin, Fin, et al.
Pubblicazione: (2026)
di: Amin, Fin, et al.
Pubblicazione: (2026)
Improving through Interaction: Searching Behavioral Representation Spaces with CMA-ES-IG
di: Dennler, Nathaniel, et al.
Pubblicazione: (2026)
di: Dennler, Nathaniel, et al.
Pubblicazione: (2026)
Robots That Know What to Ask: Recovering Misaligned Rewards through Targeted Explanations
di: Merker, Helena, et al.
Pubblicazione: (2026)
di: Merker, Helena, et al.
Pubblicazione: (2026)
Improving User Experience in Preference-Based Optimization of Reward Functions for Assistive Robots
di: Dennler, Nathaniel, et al.
Pubblicazione: (2024)
di: Dennler, Nathaniel, et al.
Pubblicazione: (2024)
Goal Inference from Open-Ended Dialog
di: Ma, Rachel, et al.
Pubblicazione: (2024)
di: Ma, Rachel, et al.
Pubblicazione: (2024)
Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
di: Liang, Anthony, et al.
Pubblicazione: (2026)
di: Liang, Anthony, et al.
Pubblicazione: (2026)
Preference-Conditioned Language-Guided Abstraction
di: Peng, Andi, et al.
Pubblicazione: (2024)
di: Peng, Andi, et al.
Pubblicazione: (2024)
Singing the Body Electric: The Impact of Robot Embodiment on User Expectations
di: Dennler, Nathaniel, et al.
Pubblicazione: (2024)
di: Dennler, Nathaniel, et al.
Pubblicazione: (2024)
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
di: Ma, Rachel, et al.
Pubblicazione: (2025)
di: Ma, Rachel, et al.
Pubblicazione: (2025)
IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
di: Jiang, Anqing, et al.
Pubblicazione: (2025)
di: Jiang, Anqing, et al.
Pubblicazione: (2025)
Contrastive Learning from Exploratory Actions: Leveraging Natural Interactions for Preference Elicitation
di: Dennler, Nathaniel, et al.
Pubblicazione: (2025)
di: Dennler, Nathaniel, et al.
Pubblicazione: (2025)
Aligning Robot and Human Representations
di: Bobu, Andreea, et al.
Pubblicazione: (2023)
di: Bobu, Andreea, et al.
Pubblicazione: (2023)
IRL-DAL: Safe and Adaptive Trajectory Planning for Autonomous Driving via Energy-Guided Diffusion Models
di: Miangoleh, Seyed Ahmad Hosseini, et al.
Pubblicazione: (2026)
di: Miangoleh, Seyed Ahmad Hosseini, et al.
Pubblicazione: (2026)
Automation from the Worker's Perspective
di: Armstrong, Ben, et al.
Pubblicazione: (2024)
di: Armstrong, Ben, et al.
Pubblicazione: (2024)
Visual IRL for Human-Like Robotic Manipulation
di: Asali, Ehsan, et al.
Pubblicazione: (2024)
di: Asali, Ehsan, et al.
Pubblicazione: (2024)
VernaCopter: Disambiguated Natural-Language-Driven Robot via Formal Specifications
di: van de Laar, Teun, et al.
Pubblicazione: (2024)
di: van de Laar, Teun, et al.
Pubblicazione: (2024)
Position: Olfaction Standardization is Essential for the Advancement of Embodied Artificial Intelligence
di: France, Kordel K., et al.
Pubblicazione: (2025)
di: France, Kordel K., et al.
Pubblicazione: (2025)
TreeIRL: Safe Urban Driving with Tree Search and Inverse Reinforcement Learning
di: Tomov, Momchil S., et al.
Pubblicazione: (2025)
di: Tomov, Momchil S., et al.
Pubblicazione: (2025)
Representation Alignment from Human Feedback for Cross-Embodiment Reward Learning from Mixed-Quality Demonstrations
di: Mattson, Connor, et al.
Pubblicazione: (2024)
di: Mattson, Connor, et al.
Pubblicazione: (2024)
From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models
di: Gumbsch, Christian, et al.
Pubblicazione: (2026)
di: Gumbsch, Christian, et al.
Pubblicazione: (2026)
Reward Learning from Suboptimal Demonstrations with Applications in Surgical Electrocautery
di: Karimi, Zohre, et al.
Pubblicazione: (2024)
di: Karimi, Zohre, et al.
Pubblicazione: (2024)
PRISM: Projection-based Reward Integration for Scene-Aware Real-to-Sim-to-Real Transfer with Few Demonstrations
di: Sun, Haowen, et al.
Pubblicazione: (2025)
di: Sun, Haowen, et al.
Pubblicazione: (2025)
MotIF: Motion Instruction Fine-tuning
di: Hwang, Minyoung, et al.
Pubblicazione: (2024)
di: Hwang, Minyoung, et al.
Pubblicazione: (2024)
Sample-Efficient Reinforcement Learning with Symmetry-Guided Demonstrations for Robotic Manipulation
di: Enayati, Amir M. Soufi, et al.
Pubblicazione: (2023)
di: Enayati, Amir M. Soufi, et al.
Pubblicazione: (2023)
Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment
di: Li, Chenliang, et al.
Pubblicazione: (2024)
di: Li, Chenliang, et al.
Pubblicazione: (2024)
CLARA: Classifying and Disambiguating User Commands for Reliable Interactive Robotic Agents
di: Park, Jeongeun, et al.
Pubblicazione: (2023)
di: Park, Jeongeun, et al.
Pubblicazione: (2023)
Reduced-Order Model-Guided Reinforcement Learning for Demonstration-Free Humanoid Locomotion
di: Liu, Shuai, et al.
Pubblicazione: (2025)
di: Liu, Shuai, et al.
Pubblicazione: (2025)
Evaluating and Personalizing User-Perceived Quality of Text-to-Speech Voices for Delivering Mindfulness Meditation with Different Physical Embodiments
di: Shi, Zhonghao, et al.
Pubblicazione: (2024)
di: Shi, Zhonghao, et al.
Pubblicazione: (2024)
Integrating Disambiguation and User Preferences into Large Language Models for Robot Motion Planning
di: Abugurain, Mohammed, et al.
Pubblicazione: (2024)
di: Abugurain, Mohammed, et al.
Pubblicazione: (2024)
Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models
di: Wu, Yanru, et al.
Pubblicazione: (2026)
di: Wu, Yanru, et al.
Pubblicazione: (2026)
Multi-Agent Inverse Q-Learning from Demonstrations
di: Haynam, Nathaniel, et al.
Pubblicazione: (2025)
di: Haynam, Nathaniel, et al.
Pubblicazione: (2025)
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
di: Hwang, Minjune, et al.
Pubblicazione: (2026)
di: Hwang, Minjune, et al.
Pubblicazione: (2026)
Demonstrating the Octopi-1.5 Visual-Tactile-Language Model
di: Yu, Samson, et al.
Pubblicazione: (2025)
di: Yu, Samson, et al.
Pubblicazione: (2025)
PROGRESSOR: A Perceptually Guided Reward Estimator with Self-Supervised Online Refinement
di: Ayalew, Tewodros, et al.
Pubblicazione: (2024)
di: Ayalew, Tewodros, et al.
Pubblicazione: (2024)
Human-Guided Harm Recovery for Computer Use Agents
di: Li, Christy, et al.
Pubblicazione: (2026)
di: Li, Christy, et al.
Pubblicazione: (2026)
Hierarchical Vision Language Action Model Using Success and Failure Demonstrations
di: Park, Jeongeun, et al.
Pubblicazione: (2025)
di: Park, Jeongeun, et al.
Pubblicazione: (2025)
Text2Touch: Tactile In-Hand Manipulation with LLM-Designed Reward Functions
di: Field, Harrison, et al.
Pubblicazione: (2025)
di: Field, Harrison, et al.
Pubblicazione: (2025)
LEMMo-Plan: LLM-Enhanced Learning from Multi-Modal Demonstration for Planning Sequential Contact-Rich Manipulation Tasks
di: Chen, Kejia, et al.
Pubblicazione: (2024)
di: Chen, Kejia, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Learning Contextually-Adaptive Rewards via Calibrated Features
di: Forsey-Smerek, Alexandra, et al.
Pubblicazione: (2025) -
QuickLAP: Quick Language-Action Preference Learning for Semi-Autonomous Agents
di: Nader, Jordan Abi, et al.
Pubblicazione: (2025) -
GIFT: Generalizing Intent for Flexible Test-Time Rewards
di: Amin, Fin, et al.
Pubblicazione: (2026) -
Improving through Interaction: Searching Behavioral Representation Spaces with CMA-ES-IG
di: Dennler, Nathaniel, et al.
Pubblicazione: (2026) -
Robots That Know What to Ask: Recovering Misaligned Rewards through Targeted Explanations
di: Merker, Helena, et al.
Pubblicazione: (2026)