Advancing Natural Language Formalization to First Order Logic with Fine-tuned LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Vossel, Felix, Mossakowski, Till, Gehrke, Björn
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912737351696384
author Vossel, Felix
Mossakowski, Till
Gehrke, Björn
author_facet Vossel, Felix
Mossakowski, Till
Gehrke, Björn
contents Automating the translation of natural language to first-order logic (FOL) is crucial for knowledge representation and formal methods, yet remains challenging. We present a systematic evaluation of fine-tuned LLMs for this task, comparing architectures (encoder-decoder vs. decoder-only) and training strategies. Using the MALLS and Willow datasets, we explore techniques like vocabulary extension, predicate conditioning, and multilingual training, introducing metrics for exact match, logical equivalence, and predicate alignment. Our fine-tuned Flan-T5-XXL achieves 70% accuracy with predicate lists, outperforming GPT-4o and even the DeepSeek-R1-0528 model with CoT reasoning ability as well as symbolic systems like ccg2lambda. Key findings show: (1) predicate availability boosts performance by 15-20%, (2) T5 models surpass larger decoder-only LLMs, and (3) models generalize to unseen logical arguments (FOLIO dataset) without specific training. While structural logic translation proves robust, predicate extraction emerges as the main bottleneck.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22338
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Advancing Natural Language Formalization to First Order Logic with Fine-tuned LLMs
Vossel, Felix
Mossakowski, Till
Gehrke, Björn
Computation and Language
Artificial Intelligence
03B10
I.2.7; I.2.3
Automating the translation of natural language to first-order logic (FOL) is crucial for knowledge representation and formal methods, yet remains challenging. We present a systematic evaluation of fine-tuned LLMs for this task, comparing architectures (encoder-decoder vs. decoder-only) and training strategies. Using the MALLS and Willow datasets, we explore techniques like vocabulary extension, predicate conditioning, and multilingual training, introducing metrics for exact match, logical equivalence, and predicate alignment. Our fine-tuned Flan-T5-XXL achieves 70% accuracy with predicate lists, outperforming GPT-4o and even the DeepSeek-R1-0528 model with CoT reasoning ability as well as symbolic systems like ccg2lambda. Key findings show: (1) predicate availability boosts performance by 15-20%, (2) T5 models surpass larger decoder-only LLMs, and (3) models generalize to unseen logical arguments (FOLIO dataset) without specific training. While structural logic translation proves robust, predicate extraction emerges as the main bottleneck.
title Advancing Natural Language Formalization to First Order Logic with Fine-tuned LLMs
topic Computation and Language
Artificial Intelligence
03B10
I.2.7; I.2.3
url https://arxiv.org/abs/2509.22338