Pipeline for Verifying LLM-Generated Mathematical Solutions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sazonova, Varvara, Shmelkin, Dmitri, Kikot, Stanislav, Motolygin, Vasily
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914347136057344
author Sazonova, Varvara
Shmelkin, Dmitri
Kikot, Stanislav
Motolygin, Vasily
author_facet Sazonova, Varvara
Shmelkin, Dmitri
Kikot, Stanislav
Motolygin, Vasily
contents With the growing popularity of Large Reasoning Models and their results in solving mathematical problems, it becomes crucial to measure their capabilities. We introduce a pipeline for both automatic and interactive verification as a more accurate alternative to only checking the answer which is currently the most popular approach for benchmarks. The pipeline can also be used as a generator of correct solutions both in formal and informal languages. 3 AI agents, which can be chosen for the benchmark accordingly, are included in the structure. The key idea is the use of prompts to obtain the solution in the specific form which allows for easier verification using proof assistants and possible use of small models ($\le 8B$). Experiments on several datasets suggest low probability of False Positives. The open-source implementation with instructions on setting up a server is available at https://github.com/LogicEnj/lean4_verification_pipeline.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20770
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Pipeline for Verifying LLM-Generated Mathematical Solutions
Sazonova, Varvara
Shmelkin, Dmitri
Kikot, Stanislav
Motolygin, Vasily
Artificial Intelligence
With the growing popularity of Large Reasoning Models and their results in solving mathematical problems, it becomes crucial to measure their capabilities. We introduce a pipeline for both automatic and interactive verification as a more accurate alternative to only checking the answer which is currently the most popular approach for benchmarks. The pipeline can also be used as a generator of correct solutions both in formal and informal languages. 3 AI agents, which can be chosen for the benchmark accordingly, are included in the structure. The key idea is the use of prompts to obtain the solution in the specific form which allows for easier verification using proof assistants and possible use of small models ($\le 8B$). Experiments on several datasets suggest low probability of False Positives. The open-source implementation with instructions on setting up a server is available at https://github.com/LogicEnj/lean4_verification_pipeline.
title Pipeline for Verifying LLM-Generated Mathematical Solutions
topic Artificial Intelligence
url https://arxiv.org/abs/2602.20770