Math Natural Language Inference: this should be easy!

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: de Paiva, Valeria, Gao, Qiyue, Hu, Hai, Kovalev, Pavel, Liu, Yikang, Moss, Lawrence S., Qian, Zhiheng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908473344655360
author de Paiva, Valeria
Gao, Qiyue
Hu, Hai
Kovalev, Pavel
Liu, Yikang
Moss, Lawrence S.
Qian, Zhiheng
author_facet de Paiva, Valeria
Gao, Qiyue
Hu, Hai
Kovalev, Pavel
Liu, Yikang
Moss, Lawrence S.
Qian, Zhiheng
contents We ask whether contemporary LLMs are able to perform natural language inference (NLI) tasks on mathematical texts. We call this the Math NLI problem. We construct a corpus of Math NLI pairs whose premises are from extant mathematical text and whose hypotheses and gold labels were provided by people with experience in both research-level mathematics and also in the NLI field. We also investigate the quality of corpora using the same premises but whose hypotheses are provided by LLMs themselves. We not only investigate the performance but also the inter-group consistency of the diverse group of LLMs. We have both positive and negative findings. Among our positive findings: in some settings, using a majority vote of LLMs is approximately equivalent to using human-labeled data in the Math NLI area. On the negative side: LLMs still struggle with mathematical language. They occasionally fail at even basic inferences. Current models are not as prone to hypothesis-only "inference" in our data the way the previous generation had been. In addition to our findings, we also provide our corpora as data to support future work on Math NLI.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23063
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Math Natural Language Inference: this should be easy!
de Paiva, Valeria
Gao, Qiyue
Hu, Hai
Kovalev, Pavel
Liu, Yikang
Moss, Lawrence S.
Qian, Zhiheng
Computation and Language
68T50
I.2.7
We ask whether contemporary LLMs are able to perform natural language inference (NLI) tasks on mathematical texts. We call this the Math NLI problem. We construct a corpus of Math NLI pairs whose premises are from extant mathematical text and whose hypotheses and gold labels were provided by people with experience in both research-level mathematics and also in the NLI field. We also investigate the quality of corpora using the same premises but whose hypotheses are provided by LLMs themselves. We not only investigate the performance but also the inter-group consistency of the diverse group of LLMs. We have both positive and negative findings. Among our positive findings: in some settings, using a majority vote of LLMs is approximately equivalent to using human-labeled data in the Math NLI area. On the negative side: LLMs still struggle with mathematical language. They occasionally fail at even basic inferences. Current models are not as prone to hypothesis-only "inference" in our data the way the previous generation had been. In addition to our findings, we also provide our corpora as data to support future work on Math NLI.
title Math Natural Language Inference: this should be easy!
topic Computation and Language
68T50
I.2.7
url https://arxiv.org/abs/2507.23063