Safe and Optimal Learning from Preferences via Weighted Temporal Logic with Applications in Robotics and Formula 1

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Karagulle, Ruya, Vasile, Cristian-Ioan, Ozay, Necmiye
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908877395591168
author Karagulle, Ruya
Vasile, Cristian-Ioan
Ozay, Necmiye
author_facet Karagulle, Ruya
Vasile, Cristian-Ioan
Ozay, Necmiye
contents Autonomous systems increasingly rely on human feedback to align their behavior, expressed as pairwise comparisons, rankings, or demonstrations. While existing methods can adapt behaviors, they often fail to guarantee safety in safety-critical domains. We propose a safety-guaranteed, optimal, and efficient approach for solving the learning problem from preferences, rankings, or demonstrations using Weighted Signal Temporal Logic (WSTL). WSTL learning problems, when implemented naively, lead to multi-linear constraints in the weights to be learned. By introducing structural pruning and log-transform procedures, we reduce the problem size and recast it as a Mixed-Integer Linear Program while preserving safety guarantees. Experiments on robotic navigation and real-world Formula 1 data demonstrate that the method captures nuanced preferences and models complex task objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08502
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Safe and Optimal Learning from Preferences via Weighted Temporal Logic with Applications in Robotics and Formula 1
Karagulle, Ruya
Vasile, Cristian-Ioan
Ozay, Necmiye
Robotics
Systems and Control
Autonomous systems increasingly rely on human feedback to align their behavior, expressed as pairwise comparisons, rankings, or demonstrations. While existing methods can adapt behaviors, they often fail to guarantee safety in safety-critical domains. We propose a safety-guaranteed, optimal, and efficient approach for solving the learning problem from preferences, rankings, or demonstrations using Weighted Signal Temporal Logic (WSTL). WSTL learning problems, when implemented naively, lead to multi-linear constraints in the weights to be learned. By introducing structural pruning and log-transform procedures, we reduce the problem size and recast it as a Mixed-Integer Linear Program while preserving safety guarantees. Experiments on robotic navigation and real-world Formula 1 data demonstrate that the method captures nuanced preferences and models complex task objectives.
title Safe and Optimal Learning from Preferences via Weighted Temporal Logic with Applications in Robotics and Formula 1
topic Robotics
Systems and Control
url https://arxiv.org/abs/2511.08502