SysMoBench: Evaluating AI on Formally Modeling Complex Real-World Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Qian, Tang, Ruize, Ma, Emilie, Hackett, Finn, He, Peiyang, Su, Yiming, Beschastnikh, Ivan, Huang, Yu, Ma, Xiaoxing, Xu, Tianyin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917226438721536
author Cheng, Qian
Tang, Ruize
Ma, Emilie
Hackett, Finn
He, Peiyang
Su, Yiming
Beschastnikh, Ivan
Huang, Yu
Ma, Xiaoxing
Xu, Tianyin
author_facet Cheng, Qian
Tang, Ruize
Ma, Emilie
Hackett, Finn
He, Peiyang
Su, Yiming
Beschastnikh, Ivan
Huang, Yu
Ma, Xiaoxing
Xu, Tianyin
contents Formal models are essential to specifying large, complex computer systems and verifying their correctness, but are notoriously expensive to write and maintain. Recent advances in generative AI show promise in generating certain forms of specifications. However, existing work mostly targets small code, not complete systems. It is unclear whether AI can deal with realistic system artifacts, as this requires abstracting their complex behavioral properties into formal models. We present SysMoBench, a benchmark that evaluates AI's ability to formally model large, complex systems. We focus on concurrent and distributed systems, which are keystones of today's critical computing infrastructures, encompassing operating systems and cloud infrastructure. We use TLA+, the de facto specification language for concurrent and distributed systems, though the benchmark can be extended to other specification languages. We address the primary challenge of evaluating AI-generated models by automating metrics like syntactic and runtime correctness, conformance to system code, and invariant correctness. SysMoBench currently includes eleven diverse system artifacts: the Raft implementation of Etcd and Redis, the leader election of ZooKeeper, the Spinlock, Mutex, and Ringbuffer in Asterinas OS, etc., with more being added. SysMoBench enables us to understand the capabilities and limitations of today's LLMs and agents, putting tools in this area on a firm footing and opening up promising new research directions.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23130
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SysMoBench: Evaluating AI on Formally Modeling Complex Real-World Systems
Cheng, Qian
Tang, Ruize
Ma, Emilie
Hackett, Finn
He, Peiyang
Su, Yiming
Beschastnikh, Ivan
Huang, Yu
Ma, Xiaoxing
Xu, Tianyin
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Software Engineering
Formal models are essential to specifying large, complex computer systems and verifying their correctness, but are notoriously expensive to write and maintain. Recent advances in generative AI show promise in generating certain forms of specifications. However, existing work mostly targets small code, not complete systems. It is unclear whether AI can deal with realistic system artifacts, as this requires abstracting their complex behavioral properties into formal models. We present SysMoBench, a benchmark that evaluates AI's ability to formally model large, complex systems. We focus on concurrent and distributed systems, which are keystones of today's critical computing infrastructures, encompassing operating systems and cloud infrastructure. We use TLA+, the de facto specification language for concurrent and distributed systems, though the benchmark can be extended to other specification languages. We address the primary challenge of evaluating AI-generated models by automating metrics like syntactic and runtime correctness, conformance to system code, and invariant correctness. SysMoBench currently includes eleven diverse system artifacts: the Raft implementation of Etcd and Redis, the leader election of ZooKeeper, the Spinlock, Mutex, and Ringbuffer in Asterinas OS, etc., with more being added. SysMoBench enables us to understand the capabilities and limitations of today's LLMs and agents, putting tools in this area on a firm footing and opening up promising new research directions.
title SysMoBench: Evaluating AI on Formally Modeling Complex Real-World Systems
topic Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Software Engineering
url https://arxiv.org/abs/2509.23130