Expressive Temporal Specifications for Reward Monitoring

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Adalat, Omar, Belardinelli, Francesco
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912792754257920
author Adalat, Omar
Belardinelli, Francesco
author_facet Adalat, Omar
Belardinelli, Francesco
contents Specifying informative and dense reward functions remains a pivotal challenge in Reinforcement Learning, as it directly affects the efficiency of agent training. In this work, we harness the expressive power of quantitative Linear Temporal Logic on finite traces (($\text{LTL}_f[\mathcal{F}]$)) to synthesize reward monitors that generate a dense stream of rewards for runtime-observable state trajectories. By providing nuanced feedback during training, these monitors guide agents toward optimal behaviour and help mitigate the well-known issue of sparse rewards under long-horizon decision making, which arises under the Boolean semantics dominating the current literature. Our framework is algorithm-agnostic and only relies on a state labelling function, and naturally accommodates specifying non-Markovian properties. Empirical results show that our quantitative monitors consistently subsume and, depending on the environment, outperform Boolean monitors in maximizing a quantitative measure of task completion and in reducing convergence time.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12808
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Expressive Temporal Specifications for Reward Monitoring
Adalat, Omar
Belardinelli, Francesco
Machine Learning
Artificial Intelligence
Logic in Computer Science
Specifying informative and dense reward functions remains a pivotal challenge in Reinforcement Learning, as it directly affects the efficiency of agent training. In this work, we harness the expressive power of quantitative Linear Temporal Logic on finite traces (($\text{LTL}_f[\mathcal{F}]$)) to synthesize reward monitors that generate a dense stream of rewards for runtime-observable state trajectories. By providing nuanced feedback during training, these monitors guide agents toward optimal behaviour and help mitigate the well-known issue of sparse rewards under long-horizon decision making, which arises under the Boolean semantics dominating the current literature. Our framework is algorithm-agnostic and only relies on a state labelling function, and naturally accommodates specifying non-Markovian properties. Empirical results show that our quantitative monitors consistently subsume and, depending on the environment, outperform Boolean monitors in maximizing a quantitative measure of task completion and in reducing convergence time.
title Expressive Temporal Specifications for Reward Monitoring
topic Machine Learning
Artificial Intelligence
Logic in Computer Science
url https://arxiv.org/abs/2511.12808