Saved in:
Bibliographic Details
Main Authors: Mehrdad, Sarmad, Meduri, Avadesh, Righetti, Ludovic
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.08619
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909609011183616
author Mehrdad, Sarmad
Meduri, Avadesh
Righetti, Ludovic
author_facet Mehrdad, Sarmad
Meduri, Avadesh
Righetti, Ludovic
contents We present an iterative inverse reinforcement learning algorithm to infer optimal cost functions in continuous spaces. Based on a popular maximum entropy criteria, our approach iteratively finds a weight improvement step and proposes a method to find an appropriate step size that ensures learned cost function features remain similar to the demonstrated trajectory features. In contrast to similar approaches, our algorithm can individually tune the effectiveness of each observation for the partition function and does not need a large sample set, enabling faster learning. We generate sample trajectories by solving an optimal control problem instead of random sampling, leading to more informative trajectories. The performance of our method is compared to two state of the art algorithms to demonstrate its benefits in several simulated environments.
format Preprint
id arxiv_https___arxiv_org_abs_2505_08619
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cost Function Estimation Using Inverse Reinforcement Learning with Minimal Observations
Mehrdad, Sarmad
Meduri, Avadesh
Righetti, Ludovic
Machine Learning
Robotics
We present an iterative inverse reinforcement learning algorithm to infer optimal cost functions in continuous spaces. Based on a popular maximum entropy criteria, our approach iteratively finds a weight improvement step and proposes a method to find an appropriate step size that ensures learned cost function features remain similar to the demonstrated trajectory features. In contrast to similar approaches, our algorithm can individually tune the effectiveness of each observation for the partition function and does not need a large sample set, enabling faster learning. We generate sample trajectories by solving an optimal control problem instead of random sampling, leading to more informative trajectories. The performance of our method is compared to two state of the art algorithms to demonstrate its benefits in several simulated environments.
title Cost Function Estimation Using Inverse Reinforcement Learning with Minimal Observations
topic Machine Learning
Robotics
url https://arxiv.org/abs/2505.08619