Value Functions for Temporal Logic: Optimal Policies and Safety Filters
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866914526019977216 |
|---|---|
| author | So, Oswin Sharpless, William Herbert, Sylvia Fan, Chuchu |
| author_facet | So, Oswin Sharpless, William Herbert, Sylvia Fan, Chuchu |
| contents | While Bellman equations for basic reach, avoid, and reach-avoid problems are well studied, the relationship between value optimality and policy optimality becomes subtle in the undiscounted infinite-horizon setting, particularly for more complicated tasks. Greedily maximizing the Q-function can produce policies that indefinitely defer task completion for reach-avoid problems, or equivalently, Until specifications, even when the value function is optimal. Building upon recent results decomposing the value function for temporal logic (TL) into a graph of constituent value functions, we construct non-Markovian policies based on state history that avoid this pathology and prove their optimality with respect to the quantitative robustness score for nested Until, Globally, and Globally-Until specifications. We further show how the Q function can serve as a safety filter for complex TL specifications, extending prior results beyond simple avoid or reach-avoid tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_01051 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Value Functions for Temporal Logic: Optimal Policies and Safety Filters So, Oswin Sharpless, William Herbert, Sylvia Fan, Chuchu Robotics Artificial Intelligence Machine Learning Logic in Computer Science Optimization and Control While Bellman equations for basic reach, avoid, and reach-avoid problems are well studied, the relationship between value optimality and policy optimality becomes subtle in the undiscounted infinite-horizon setting, particularly for more complicated tasks. Greedily maximizing the Q-function can produce policies that indefinitely defer task completion for reach-avoid problems, or equivalently, Until specifications, even when the value function is optimal. Building upon recent results decomposing the value function for temporal logic (TL) into a graph of constituent value functions, we construct non-Markovian policies based on state history that avoid this pathology and prove their optimality with respect to the quantitative robustness score for nested Until, Globally, and Globally-Until specifications. We further show how the Q function can serve as a safety filter for complex TL specifications, extending prior results beyond simple avoid or reach-avoid tasks. |
| title | Value Functions for Temporal Logic: Optimal Policies and Safety Filters |
| topic | Robotics Artificial Intelligence Machine Learning Logic in Computer Science Optimization and Control |
| url | https://arxiv.org/abs/2605.01051 |