Can Large Language Models Autoformalize Kinematics?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kabra, Aditi, Laurent, Jonathan, Bharadwaj, Sagar, Martins, Ruben, Mitsch, Stefan, Platzer, André
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918155378491392
author Kabra, Aditi
Laurent, Jonathan
Bharadwaj, Sagar
Martins, Ruben
Mitsch, Stefan
Platzer, André
author_facet Kabra, Aditi
Laurent, Jonathan
Bharadwaj, Sagar
Martins, Ruben
Mitsch, Stefan
Platzer, André
contents Autonomous cyber-physical systems like robots and self-driving cars could greatly benefit from using formal methods to reason reliably about their control decisions. However, before a problem can be solved it needs to be stated. This requires writing a formal physics model of the cyber-physical system, which is a complex task that traditionally requires human expertise and becomes a bottleneck. This paper experimentally studies whether Large Language Models (LLMs) can automate the formalization process. A 20 problem benchmark suite is designed drawing from undergraduate level physics kinematics problems. In each problem, the LLM is provided with a natural language description of the objects' motion and must produce a model in differential game logic (dGL). The model is (1) syntax checked and iteratively refined based on parser feedback, and (2) semantically evaluated by checking whether symbolically executing the dGL formula recovers the solution to the original physics problem. A success rate of 70% (best over 5 samples) is achieved. We analyze failing cases, identifying directions for future improvement. This provides a first quantitative baseline for LLM-based autoformalization from natural language to a hybrid games logic with continuous dynamics.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21840
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Large Language Models Autoformalize Kinematics?
Kabra, Aditi
Laurent, Jonathan
Bharadwaj, Sagar
Martins, Ruben
Mitsch, Stefan
Platzer, André
Logic in Computer Science
Artificial Intelligence
Autonomous cyber-physical systems like robots and self-driving cars could greatly benefit from using formal methods to reason reliably about their control decisions. However, before a problem can be solved it needs to be stated. This requires writing a formal physics model of the cyber-physical system, which is a complex task that traditionally requires human expertise and becomes a bottleneck. This paper experimentally studies whether Large Language Models (LLMs) can automate the formalization process. A 20 problem benchmark suite is designed drawing from undergraduate level physics kinematics problems. In each problem, the LLM is provided with a natural language description of the objects' motion and must produce a model in differential game logic (dGL). The model is (1) syntax checked and iteratively refined based on parser feedback, and (2) semantically evaluated by checking whether symbolically executing the dGL formula recovers the solution to the original physics problem. A success rate of 70% (best over 5 samples) is achieved. We analyze failing cases, identifying directions for future improvement. This provides a first quantitative baseline for LLM-based autoformalization from natural language to a hybrid games logic with continuous dynamics.
title Can Large Language Models Autoformalize Kinematics?
topic Logic in Computer Science
Artificial Intelligence
url https://arxiv.org/abs/2509.21840