Towards Autoformalization of LLM-generated Outputs for Requirement Verification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gupte, Mihir, S, Ramesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911266542452736
author Gupte, Mihir
S, Ramesh
author_facet Gupte, Mihir
S, Ramesh
contents Autoformalization, the process of translating informal statements into formal logic, has gained renewed interest with the emergence of powerful Large Language Models (LLMs). While LLMs show promise in generating structured outputs from natural language (NL), such as Gherkin Scenarios from NL feature requirements, there's currently no formal method to verify if these outputs are accurate. This paper takes a preliminary step toward addressing this gap by exploring the use of a simple LLM-based autoformalizer to verify LLM-generated outputs against a small set of natural language requirements. We conducted two distinct experiments. In the first one, the autoformalizer successfully identified that two differently-worded NL requirements were logically equivalent, demonstrating the pipeline's potential for consistency checks. In the second, the autoformalizer was used to identify a logical inconsistency between a given NL requirement and an LLM-generated output, highlighting its utility as a formal verification tool. Our findings, while limited, suggest that autoformalization holds significant potential for ensuring the fidelity and logical consistency of LLM-generated outputs, laying a crucial foundation for future, more extensive studies into this novel application.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11829
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Autoformalization of LLM-generated Outputs for Requirement Verification
Gupte, Mihir
S, Ramesh
Computation and Language
Artificial Intelligence
Formal Languages and Automata Theory
Logic in Computer Science
Autoformalization, the process of translating informal statements into formal logic, has gained renewed interest with the emergence of powerful Large Language Models (LLMs). While LLMs show promise in generating structured outputs from natural language (NL), such as Gherkin Scenarios from NL feature requirements, there's currently no formal method to verify if these outputs are accurate. This paper takes a preliminary step toward addressing this gap by exploring the use of a simple LLM-based autoformalizer to verify LLM-generated outputs against a small set of natural language requirements. We conducted two distinct experiments. In the first one, the autoformalizer successfully identified that two differently-worded NL requirements were logically equivalent, demonstrating the pipeline's potential for consistency checks. In the second, the autoformalizer was used to identify a logical inconsistency between a given NL requirement and an LLM-generated output, highlighting its utility as a formal verification tool. Our findings, while limited, suggest that autoformalization holds significant potential for ensuring the fidelity and logical consistency of LLM-generated outputs, laying a crucial foundation for future, more extensive studies into this novel application.
title Towards Autoformalization of LLM-generated Outputs for Requirement Verification
topic Computation and Language
Artificial Intelligence
Formal Languages and Automata Theory
Logic in Computer Science
url https://arxiv.org/abs/2511.11829