Automated Formalization via Conceptual Retrieval-Augmented LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lu, Wangyue, Du, Lun, Li, Sirui, Weng, Ke, Sun, Haozhe, Liu, Hengyu, Yu, Minghe, Zhang, Tiancheng, Yu, Ge
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915879343620096
author Lu, Wangyue
Du, Lun
Li, Sirui
Weng, Ke
Sun, Haozhe
Liu, Hengyu
Yu, Minghe
Zhang, Tiancheng
Yu, Ge
author_facet Lu, Wangyue
Du, Lun
Li, Sirui
Weng, Ke
Sun, Haozhe
Liu, Hengyu
Yu, Minghe
Zhang, Tiancheng
Yu, Ge
contents Interactive theorem provers (ITPs) require manual formalization, which is labor-intensive and demands expert knowledge. While automated formalization offers a potential solution, it faces two major challenges: model hallucination (e.g., undefined predicates, symbol misuse, and version incompatibility) and the semantic gap caused by ambiguous or missing premises in natural language descriptions. To address these issues, we propose CRAMF, a Concept-driven Retrieval-Augmented Mathematical Formalization framework. CRAMF enhances LLM-based autoformalization by retrieving formal definitions of core mathematical concepts, providing contextual grounding during code generation. However, applying retrieval-augmented generation (RAG) in this setting is non-trivial due to the lack of structured knowledge bases, the polymorphic nature of mathematical concepts, and the high precision required in formal retrieval. We introduce a framework for automatically constructing a concept-definition knowledge base from Mathlib4, the standard mathematical library for the Lean 4 theorem prover, indexing over 26,000 formal definitions and 1,000+ core mathematical concepts. To address conceptual polymorphism, we propose contextual query augmentation with domain- and application-level signals. In addition, we design a dual-channel hybrid retrieval strategy with reranking to ensure accurate and relevant definition retrieval. Experiments on miniF2F, ProofNet, and our newly proposed AdvancedMath benchmark show that CRAMF can be seamlessly integrated into LLM-based autoformalizers, yielding consistent improvements in translation accuracy, achieving up to 62.1% and an average of 29.9% relative improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06931
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Automated Formalization via Conceptual Retrieval-Augmented LLMs
Lu, Wangyue
Du, Lun
Li, Sirui
Weng, Ke
Sun, Haozhe
Liu, Hengyu
Yu, Minghe
Zhang, Tiancheng
Yu, Ge
Artificial Intelligence
Machine Learning
Interactive theorem provers (ITPs) require manual formalization, which is labor-intensive and demands expert knowledge. While automated formalization offers a potential solution, it faces two major challenges: model hallucination (e.g., undefined predicates, symbol misuse, and version incompatibility) and the semantic gap caused by ambiguous or missing premises in natural language descriptions. To address these issues, we propose CRAMF, a Concept-driven Retrieval-Augmented Mathematical Formalization framework. CRAMF enhances LLM-based autoformalization by retrieving formal definitions of core mathematical concepts, providing contextual grounding during code generation. However, applying retrieval-augmented generation (RAG) in this setting is non-trivial due to the lack of structured knowledge bases, the polymorphic nature of mathematical concepts, and the high precision required in formal retrieval. We introduce a framework for automatically constructing a concept-definition knowledge base from Mathlib4, the standard mathematical library for the Lean 4 theorem prover, indexing over 26,000 formal definitions and 1,000+ core mathematical concepts. To address conceptual polymorphism, we propose contextual query augmentation with domain- and application-level signals. In addition, we design a dual-channel hybrid retrieval strategy with reranking to ensure accurate and relevant definition retrieval. Experiments on miniF2F, ProofNet, and our newly proposed AdvancedMath benchmark show that CRAMF can be seamlessly integrated into LLM-based autoformalizers, yielding consistent improvements in translation accuracy, achieving up to 62.1% and an average of 29.9% relative improvement.
title Automated Formalization via Conceptual Retrieval-Augmented LLMs
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.06931