LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Xuhan, Shen, Qingning, Hu, Yan, Gao, Anningzhe, Wang, Benyou
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916615600209920
author Huang, Xuhan
Shen, Qingning
Hu, Yan
Gao, Anningzhe
Wang, Benyou
author_facet Huang, Xuhan
Shen, Qingning
Hu, Yan
Gao, Anningzhe
Wang, Benyou
contents Large Language Models (LLMs) have demonstrated strong performance across various natural language processing tasks, yet their proficiency in mathematical reasoning remains a key challenge. Addressing the gap between natural and mathematical language requires advanced reasoning capabilities, approaching those of Artificial General Intelligence (AGI). However, the evaluation remains challenging, as perfectly representing reality is inherently elusive, and traditional methods like manual or direct comparison of mathematical statements (Ramamonjison et al., 2023) are insufficient for assessing true modeling ability. We propose a process-oriented framework to evaluate LLMs' ability to construct mathematical models, using solvers to compare outputs with ground truth. Introducing Mamo, a benchmark with 1,209 questions covering ordinary differential equations, linear programming, and mixed-integer linear programming, we enable automatic evaluation of modeling accuracy. The results show that existing LLMs struggle with complex mathematical modeling tasks, with larger models demonstrating superior performance, while open-source models remain competitive in simpler cases but still fall short of proprietary models in more challenging problems.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13144
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages
Huang, Xuhan
Shen, Qingning
Hu, Yan
Gao, Anningzhe
Wang, Benyou
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) have demonstrated strong performance across various natural language processing tasks, yet their proficiency in mathematical reasoning remains a key challenge. Addressing the gap between natural and mathematical language requires advanced reasoning capabilities, approaching those of Artificial General Intelligence (AGI). However, the evaluation remains challenging, as perfectly representing reality is inherently elusive, and traditional methods like manual or direct comparison of mathematical statements (Ramamonjison et al., 2023) are insufficient for assessing true modeling ability. We propose a process-oriented framework to evaluate LLMs' ability to construct mathematical models, using solvers to compare outputs with ground truth. Introducing Mamo, a benchmark with 1,209 questions covering ordinary differential equations, linear programming, and mixed-integer linear programming, we enable automatic evaluation of modeling accuracy. The results show that existing LLMs struggle with complex mathematical modeling tasks, with larger models demonstrating superior performance, while open-source models remain competitive in simpler cases but still fall short of proprietary models in more challenging problems.
title LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2405.13144