CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Zhengyang, Ye, Zihan, Huang, Chenyu, Huang, Xuhan, Li, Chengpeng, Li, Sihang, Chen, Guanhua, Yan, Ming, Wang, Zizhuo, Zha, Hongyuan, Liu, Dayiheng, Wang, Benyou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912630207152128
author Tang, Zhengyang
Ye, Zihan
Huang, Chenyu
Huang, Xuhan
Li, Chengpeng
Li, Sihang
Chen, Guanhua
Yan, Ming
Wang, Zizhuo
Zha, Hongyuan
Liu, Dayiheng
Wang, Benyou
author_facet Tang, Zhengyang
Ye, Zihan
Huang, Chenyu
Huang, Xuhan
Li, Chengpeng
Li, Sihang
Chen, Guanhua
Yan, Ming
Wang, Zizhuo
Zha, Hongyuan
Liu, Dayiheng
Wang, Benyou
contents Large Reasoning Models (LRMs) have demonstrated strong capabilities in complex multi-step reasoning, opening new opportunities for automating optimization modeling. However, existing domain adaptation methods, originally designed for earlier instruction-tuned models, often fail to exploit the advanced reasoning patterns of modern LRMs -- In particular, we show that direct fine-tuning on traditional \textit{non-reflective} datasets leads to limited gains. To fully leverage LRMs' inherent reasoning abilities, we propose \textbf{CALM} (\textit{Corrective Adaptation with Lightweight Modification}), a framework that progressively refines LRMs within their native reasoning modes for optimization modeling tasks. In CALM, an expert intervener identifies reasoning flaws and provides concise corrective hints, which the LRM incorporates to produce improved reasoning trajectories. These interventions modify fewer than 2.6\% of generated tokens, but generate high-quality data for soft adaptation through supervised fine-tuning. The adapted model is then further improved through reinforcement learning. Building on CALM, we develop \textbf{STORM} (\textit{Smart Thinking Optimization Reasoning Model}), a 4B-parameter LRM that achieves a new state-of-the-art average accuracy of 68.9\% across five popular optimization modeling benchmarks, matching the performance of a 671B LRM. These results demonstrate that dynamic, hint-based data synthesis both preserves and amplifies the native reasoning patterns of modern LRMs, offering a more effective and scalable path towards expert-level performance on challenging optimization modeling tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04204
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling
Tang, Zhengyang
Ye, Zihan
Huang, Chenyu
Huang, Xuhan
Li, Chengpeng
Li, Sihang
Chen, Guanhua
Yan, Ming
Wang, Zizhuo
Zha, Hongyuan
Liu, Dayiheng
Wang, Benyou
Computation and Language
Artificial Intelligence
Computational Engineering, Finance, and Science
Machine Learning
Large Reasoning Models (LRMs) have demonstrated strong capabilities in complex multi-step reasoning, opening new opportunities for automating optimization modeling. However, existing domain adaptation methods, originally designed for earlier instruction-tuned models, often fail to exploit the advanced reasoning patterns of modern LRMs -- In particular, we show that direct fine-tuning on traditional \textit{non-reflective} datasets leads to limited gains. To fully leverage LRMs' inherent reasoning abilities, we propose \textbf{CALM} (\textit{Corrective Adaptation with Lightweight Modification}), a framework that progressively refines LRMs within their native reasoning modes for optimization modeling tasks. In CALM, an expert intervener identifies reasoning flaws and provides concise corrective hints, which the LRM incorporates to produce improved reasoning trajectories. These interventions modify fewer than 2.6\% of generated tokens, but generate high-quality data for soft adaptation through supervised fine-tuning. The adapted model is then further improved through reinforcement learning. Building on CALM, we develop \textbf{STORM} (\textit{Smart Thinking Optimization Reasoning Model}), a 4B-parameter LRM that achieves a new state-of-the-art average accuracy of 68.9\% across five popular optimization modeling benchmarks, matching the performance of a 671B LRM. These results demonstrate that dynamic, hint-based data synthesis both preserves and amplifies the native reasoning patterns of modern LRMs, offering a more effective and scalable path towards expert-level performance on challenging optimization modeling tasks.
title CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling
topic Computation and Language
Artificial Intelligence
Computational Engineering, Finance, and Science
Machine Learning
url https://arxiv.org/abs/2510.04204