OR-Toolformer: Modeling and Solving Operations Research Problems with Tool Augmented Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jianzhang, Zhou, Jialong, Liu, Chuang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912622244265984
author Zhang, Jianzhang
Zhou, Jialong
Liu, Chuang
author_facet Zhang, Jianzhang
Zhou, Jialong
Liu, Chuang
contents Large language models (LLMs) demonstrate strong mathematical reasoning, but reliance on closed-source APIs for OR tasks raises privacy concerns, and training open-source models from scratch incurs high compute costs. We introduce OR-Toolformer, which fine-tunes Llama-3.1-8B-Instruct with a semi-automatic data synthesis pipeline that generates diverse OR problem-answer pairs and augments the model with external solvers to produce API calls. On three of four standard benchmarks, OR-Toolformer achieves up to 80.1% execution accuracy, exceeding size-matched baselines by over 4.3%. In zero-shot evaluation on two unseen OR problem types, it attains 54% average accuracy, a 21 percentage-point improvement over the strongest baseline. These findings validate the efficacy of tool-augmented fine-tuning LLMs for accurate and generalizable OR problem modeling and solving.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01253
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OR-Toolformer: Modeling and Solving Operations Research Problems with Tool Augmented Large Language Models
Zhang, Jianzhang
Zhou, Jialong
Liu, Chuang
Artificial Intelligence
Machine Learning
Large language models (LLMs) demonstrate strong mathematical reasoning, but reliance on closed-source APIs for OR tasks raises privacy concerns, and training open-source models from scratch incurs high compute costs. We introduce OR-Toolformer, which fine-tunes Llama-3.1-8B-Instruct with a semi-automatic data synthesis pipeline that generates diverse OR problem-answer pairs and augments the model with external solvers to produce API calls. On three of four standard benchmarks, OR-Toolformer achieves up to 80.1% execution accuracy, exceeding size-matched baselines by over 4.3%. In zero-shot evaluation on two unseen OR problem types, it attains 54% average accuracy, a 21 percentage-point improvement over the strongest baseline. These findings validate the efficacy of tool-augmented fine-tuning LLMs for accurate and generalizable OR problem modeling and solving.
title OR-Toolformer: Modeling and Solving Operations Research Problems with Tool Augmented Large Language Models
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.01253