Value Functions as Supermartingale Certificates

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Abate, Alessandro, Contro, Daniel, Giacobbe, Mirco, Martínez-Suñé, Agustín, Roy, Diptarko
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916066604613632
author Abate, Alessandro
Contro, Daniel
Giacobbe, Mirco
Martínez-Suñé, Agustín
Roy, Diptarko
author_facet Abate, Alessandro
Contro, Daniel
Giacobbe, Mirco
Martínez-Suñé, Agustín
Roy, Diptarko
contents Certification methods for stochastic systems provide sufficient proof rules, based on real-valued supermartingale certificates, to determine the almost-sure satisfaction of $ω$-regular properties (and therefore of linear temporal logic) over general state spaces, encompassing both countably infinite and continuous state spaces. Conversely, reinforcement learning (RL) methods for $ω$-regular tasks have received considerable attention, but they typically lack formal guarantees that the learned policy satisfies the specification, except possibly for finite state and action spaces. We bridge these two lines of research by establishing a novel theoretical connection: under an appropriate reward, the value function associated to a policy that almost surely satisfies an $ω$-regular property encodes a Streett supermartingale certificate for that specification. Our results, validated experimentally on finite Markov decision processes, hold for finite, countably infinite, and continuous state spaces, suggesting a principled route to certificate synthesis via RL.
format Preprint
id arxiv_https___arxiv_org_abs_2605_31524
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Value Functions as Supermartingale Certificates
Abate, Alessandro
Contro, Daniel
Giacobbe, Mirco
Martínez-Suñé, Agustín
Roy, Diptarko
Machine Learning
Logic in Computer Science
Certification methods for stochastic systems provide sufficient proof rules, based on real-valued supermartingale certificates, to determine the almost-sure satisfaction of $ω$-regular properties (and therefore of linear temporal logic) over general state spaces, encompassing both countably infinite and continuous state spaces. Conversely, reinforcement learning (RL) methods for $ω$-regular tasks have received considerable attention, but they typically lack formal guarantees that the learned policy satisfies the specification, except possibly for finite state and action spaces. We bridge these two lines of research by establishing a novel theoretical connection: under an appropriate reward, the value function associated to a policy that almost surely satisfies an $ω$-regular property encodes a Streett supermartingale certificate for that specification. Our results, validated experimentally on finite Markov decision processes, hold for finite, countably infinite, and continuous state spaces, suggesting a principled route to certificate synthesis via RL.
title Value Functions as Supermartingale Certificates
topic Machine Learning
Logic in Computer Science
url https://arxiv.org/abs/2605.31524