Closing the Intent-to-Behavior Gap via Fulfillment Priority Logic

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mabsout, Bassel El, Abdelgawad, Abdelrahman, Mancuso, Renato
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911247735193600
author Mabsout, Bassel El
Abdelgawad, Abdelrahman
Mancuso, Renato
author_facet Mabsout, Bassel El
Abdelgawad, Abdelrahman
Mancuso, Renato
contents Practitioners designing reinforcement learning policies face a fundamental challenge: translating intended behavioral objectives into representative reward functions. This challenge stems from behavioral intent requiring simultaneous achievement of multiple competing objectives, typically addressed through labor-intensive linear reward composition that yields brittle results. Consider the ubiquitous robotics scenario where performance maximization directly conflicts with energy conservation. Such competitive dynamics are resistant to simple linear reward combinations. In this paper, we present the concept of objective fulfillment upon which we build Fulfillment Priority Logic (FPL). FPL allows practitioners to define logical formula representing their intentions and priorities within multi-objective reinforcement learning. Our novel Balanced Policy Gradient algorithm leverages FPL specifications to achieve up to 500\% better sample efficiency compared to Soft Actor Critic. Notably, this work constitutes the first implementation of non-linear utility scalarization design, specifically for continuous control problems.
format Preprint
id arxiv_https___arxiv_org_abs_2503_05818
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Closing the Intent-to-Behavior Gap via Fulfillment Priority Logic
Mabsout, Bassel El
Abdelgawad, Abdelrahman
Mancuso, Renato
Machine Learning
Robotics
Practitioners designing reinforcement learning policies face a fundamental challenge: translating intended behavioral objectives into representative reward functions. This challenge stems from behavioral intent requiring simultaneous achievement of multiple competing objectives, typically addressed through labor-intensive linear reward composition that yields brittle results. Consider the ubiquitous robotics scenario where performance maximization directly conflicts with energy conservation. Such competitive dynamics are resistant to simple linear reward combinations. In this paper, we present the concept of objective fulfillment upon which we build Fulfillment Priority Logic (FPL). FPL allows practitioners to define logical formula representing their intentions and priorities within multi-objective reinforcement learning. Our novel Balanced Policy Gradient algorithm leverages FPL specifications to achieve up to 500\% better sample efficiency compared to Soft Actor Critic. Notably, this work constitutes the first implementation of non-linear utility scalarization design, specifically for continuous control problems.
title Closing the Intent-to-Behavior Gap via Fulfillment Priority Logic
topic Machine Learning
Robotics
url https://arxiv.org/abs/2503.05818