Value Functions for Temporal Logic: Optimal Policies and Safety Filters

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: So, Oswin, Sharpless, William, Herbert, Sylvia, Fan, Chuchu
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914526019977216
author So, Oswin
Sharpless, William
Herbert, Sylvia
Fan, Chuchu
author_facet So, Oswin
Sharpless, William
Herbert, Sylvia
Fan, Chuchu
contents While Bellman equations for basic reach, avoid, and reach-avoid problems are well studied, the relationship between value optimality and policy optimality becomes subtle in the undiscounted infinite-horizon setting, particularly for more complicated tasks. Greedily maximizing the Q-function can produce policies that indefinitely defer task completion for reach-avoid problems, or equivalently, Until specifications, even when the value function is optimal. Building upon recent results decomposing the value function for temporal logic (TL) into a graph of constituent value functions, we construct non-Markovian policies based on state history that avoid this pathology and prove their optimality with respect to the quantitative robustness score for nested Until, Globally, and Globally-Until specifications. We further show how the Q function can serve as a safety filter for complex TL specifications, extending prior results beyond simple avoid or reach-avoid tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_01051
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Value Functions for Temporal Logic: Optimal Policies and Safety Filters
So, Oswin
Sharpless, William
Herbert, Sylvia
Fan, Chuchu
Robotics
Artificial Intelligence
Machine Learning
Logic in Computer Science
Optimization and Control
While Bellman equations for basic reach, avoid, and reach-avoid problems are well studied, the relationship between value optimality and policy optimality becomes subtle in the undiscounted infinite-horizon setting, particularly for more complicated tasks. Greedily maximizing the Q-function can produce policies that indefinitely defer task completion for reach-avoid problems, or equivalently, Until specifications, even when the value function is optimal. Building upon recent results decomposing the value function for temporal logic (TL) into a graph of constituent value functions, we construct non-Markovian policies based on state history that avoid this pathology and prove their optimality with respect to the quantitative robustness score for nested Until, Globally, and Globally-Until specifications. We further show how the Q function can serve as a safety filter for complex TL specifications, extending prior results beyond simple avoid or reach-avoid tasks.
title Value Functions for Temporal Logic: Optimal Policies and Safety Filters
topic Robotics
Artificial Intelligence
Machine Learning
Logic in Computer Science
Optimization and Control
url https://arxiv.org/abs/2605.01051