ZORMS-LfD: Learning from Demonstrations with Zeroth-Order Random Matrix Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dry, Olivia, Molloy, Timothy L., Jin, Wanxin, Shames, Iman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911072235028480
author Dry, Olivia
Molloy, Timothy L.
Jin, Wanxin
Shames, Iman
author_facet Dry, Olivia
Molloy, Timothy L.
Jin, Wanxin
Shames, Iman
contents We propose Zeroth-Order Random Matrix Search for Learning from Demonstrations (ZORMS-LfD). ZORMS-LfD enables the costs, constraints, and dynamics of constrained optimal control problems, in both continuous and discrete time, to be learned from expert demonstrations without requiring smoothness of the learning-loss landscape. In contrast, existing state-of-the-art first-order methods require the existence and computation of gradients of the costs, constraints, dynamics, and learning loss with respect to states, controls and/or parameters. Most existing methods are also tailored to discrete time, with constrained problems in continuous time receiving only cursory attention. We demonstrate that ZORMS-LfD matches or surpasses the performance of state-of-the-art methods in terms of both learning loss and compute time across a variety of benchmark problems. On unconstrained continuous-time benchmark problems, ZORMS-LfD achieves similar loss performance to state-of-the-art first-order methods with an over $80$\% reduction in compute time. On constrained continuous-time benchmark problems where there is no specialized state-of-the-art method, ZORMS-LfD is shown to outperform the commonly used gradient-free Nelder-Mead optimization method.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17096
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ZORMS-LfD: Learning from Demonstrations with Zeroth-Order Random Matrix Search
Dry, Olivia
Molloy, Timothy L.
Jin, Wanxin
Shames, Iman
Machine Learning
Numerical Analysis
Systems and Control
Optimization and Control
We propose Zeroth-Order Random Matrix Search for Learning from Demonstrations (ZORMS-LfD). ZORMS-LfD enables the costs, constraints, and dynamics of constrained optimal control problems, in both continuous and discrete time, to be learned from expert demonstrations without requiring smoothness of the learning-loss landscape. In contrast, existing state-of-the-art first-order methods require the existence and computation of gradients of the costs, constraints, dynamics, and learning loss with respect to states, controls and/or parameters. Most existing methods are also tailored to discrete time, with constrained problems in continuous time receiving only cursory attention. We demonstrate that ZORMS-LfD matches or surpasses the performance of state-of-the-art methods in terms of both learning loss and compute time across a variety of benchmark problems. On unconstrained continuous-time benchmark problems, ZORMS-LfD achieves similar loss performance to state-of-the-art first-order methods with an over $80$\% reduction in compute time. On constrained continuous-time benchmark problems where there is no specialized state-of-the-art method, ZORMS-LfD is shown to outperform the commonly used gradient-free Nelder-Mead optimization method.
title ZORMS-LfD: Learning from Demonstrations with Zeroth-Order Random Matrix Search
topic Machine Learning
Numerical Analysis
Systems and Control
Optimization and Control
url https://arxiv.org/abs/2507.17096