ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qian, Cheng, Du, Hongyi, Wang, Hongru, Chen, Xiusi, Zhang, Yuji, Sil, Avirup, Zhai, Chengxiang, McKeown, Kathleen, Ji, Heng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915295715655680
author Qian, Cheng
Du, Hongyi
Wang, Hongru
Chen, Xiusi
Zhang, Yuji
Sil, Avirup
Zhai, Chengxiang
McKeown, Kathleen
Ji, Heng
author_facet Qian, Cheng
Du, Hongyi
Wang, Hongru
Chen, Xiusi
Zhang, Yuji
Sil, Avirup
Zhai, Chengxiang
McKeown, Kathleen
Ji, Heng
contents Recent progress in large language models (LLMs) has enabled substantial advances in solving mathematical problems. However, existing benchmarks often fail to reflect the complexity of real-world problems, which demand open-ended, interdisciplinary reasoning and integration of computational tools. To address this gap, we introduce ModelingBench, a novel benchmark featuring real-world-inspired, open-ended problems from math modeling competitions across diverse domains, ranging from urban traffic optimization to ecosystem resource planning. These tasks require translating natural language into formal mathematical formulations, applying appropriate tools, and producing structured, defensible reports. ModelingBench also supports multiple valid solutions, capturing the ambiguity and creativity of practical modeling. We also present ModelingAgent, a multi-agent framework that coordinates tool use, supports structured workflows, and enables iterative self-refinement to generate well-grounded, creative solutions. To evaluate outputs, we further propose ModelingJudge, an expert-in-the-loop system leveraging LLMs as domain-specialized judges assessing solutions from multiple expert perspectives. Empirical results show that ModelingAgent substantially outperforms strong baselines and often produces solutions indistinguishable from those of human experts. Together, our work provides a comprehensive framework for evaluating and advancing real-world problem-solving in open-ended, interdisciplinary modeling challenges.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15068
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges
Qian, Cheng
Du, Hongyi
Wang, Hongru
Chen, Xiusi
Zhang, Yuji
Sil, Avirup
Zhai, Chengxiang
McKeown, Kathleen
Ji, Heng
Artificial Intelligence
Computation and Language
Machine Learning
Recent progress in large language models (LLMs) has enabled substantial advances in solving mathematical problems. However, existing benchmarks often fail to reflect the complexity of real-world problems, which demand open-ended, interdisciplinary reasoning and integration of computational tools. To address this gap, we introduce ModelingBench, a novel benchmark featuring real-world-inspired, open-ended problems from math modeling competitions across diverse domains, ranging from urban traffic optimization to ecosystem resource planning. These tasks require translating natural language into formal mathematical formulations, applying appropriate tools, and producing structured, defensible reports. ModelingBench also supports multiple valid solutions, capturing the ambiguity and creativity of practical modeling. We also present ModelingAgent, a multi-agent framework that coordinates tool use, supports structured workflows, and enables iterative self-refinement to generate well-grounded, creative solutions. To evaluate outputs, we further propose ModelingJudge, an expert-in-the-loop system leveraging LLMs as domain-specialized judges assessing solutions from multiple expert perspectives. Empirical results show that ModelingAgent substantially outperforms strong baselines and often produces solutions indistinguishable from those of human experts. Together, our work provides a comprehensive framework for evaluating and advancing real-world problem-solving in open-ended, interdisciplinary modeling challenges.
title ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.15068