Reclaiming the Source of Programmatic Policies: Programmatic versus Latent Spaces

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Carvalho, Tales H., Tjhia, Kenneth, Lelis, Levi H. S.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910652328574976
author Carvalho, Tales H.
Tjhia, Kenneth
Lelis, Levi H. S.
author_facet Carvalho, Tales H.
Tjhia, Kenneth
Lelis, Levi H. S.
contents Recent works have introduced LEAPS and HPRL, systems that learn latent spaces of domain-specific languages, which are used to define programmatic policies for partially observable Markov decision processes (POMDPs). These systems induce a latent space while optimizing losses such as the behavior loss, which aim to achieve locality in program behavior, meaning that vectors close in the latent space should correspond to similarly behaving programs. In this paper, we show that the programmatic space, induced by the domain-specific language and requiring no training, presents values for the behavior loss similar to those observed in latent spaces presented in previous work. Moreover, algorithms searching in the programmatic space significantly outperform those in LEAPS and HPRL. To explain our results, we measured the "friendliness" of the two spaces to local search algorithms. We discovered that algorithms are more likely to stop at local maxima when searching in the latent space than when searching in the programmatic space. This implies that the optimization topology of the programmatic space, induced by the reward function in conjunction with the neighborhood function, is more conducive to search than that of the latent space. This result provides an explanation for the superior performance in the programmatic space.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12166
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reclaiming the Source of Programmatic Policies: Programmatic versus Latent Spaces
Carvalho, Tales H.
Tjhia, Kenneth
Lelis, Levi H. S.
Machine Learning
Artificial Intelligence
Recent works have introduced LEAPS and HPRL, systems that learn latent spaces of domain-specific languages, which are used to define programmatic policies for partially observable Markov decision processes (POMDPs). These systems induce a latent space while optimizing losses such as the behavior loss, which aim to achieve locality in program behavior, meaning that vectors close in the latent space should correspond to similarly behaving programs. In this paper, we show that the programmatic space, induced by the domain-specific language and requiring no training, presents values for the behavior loss similar to those observed in latent spaces presented in previous work. Moreover, algorithms searching in the programmatic space significantly outperform those in LEAPS and HPRL. To explain our results, we measured the "friendliness" of the two spaces to local search algorithms. We discovered that algorithms are more likely to stop at local maxima when searching in the latent space than when searching in the programmatic space. This implies that the optimization topology of the programmatic space, induced by the reward function in conjunction with the neighborhood function, is more conducive to search than that of the latent space. This result provides an explanation for the superior performance in the programmatic space.
title Reclaiming the Source of Programmatic Policies: Programmatic versus Latent Spaces
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.12166