DualSchool: How Reliable are LLMs for Optimization Education?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Klamkin, Michael, Deza, Arnaud, Cheng, Sikai, Zhao, Haoruo, Van Hentenryck, Pascal
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913862422364160
author Klamkin, Michael
Deza, Arnaud
Cheng, Sikai
Zhao, Haoruo
Van Hentenryck, Pascal
author_facet Klamkin, Michael
Deza, Arnaud
Cheng, Sikai
Zhao, Haoruo
Van Hentenryck, Pascal
contents Consider the following task taught in introductory optimization courses which addresses challenges articulated by the community at the intersection of (generative) AI and OR: generate the dual of a linear program. LLMs, being trained at web-scale, have the conversion process and many instances of Primal to Dual Conversion (P2DC) at their disposal. Students may thus reasonably expect that LLMs would perform well on the P2DC task. To assess this expectation, this paper introduces DualSchool, a comprehensive framework for generating and verifying P2DC instances. The verification procedure of DualSchool uses the Canonical Graph Edit Distance, going well beyond existing evaluation methods for optimization models, which exhibit many false positives and negatives when applied to P2DC. Experiments performed by DualSchool reveal interesting findings. Although LLMs can recite the conversion procedure accurately, state-of-the-art open LLMs fail to consistently produce correct duals. This finding holds even for the smallest two-variable instances and for derivative tasks, such as correctness, verification, and error classification. The paper also discusses the implications for educators, students, and the development of large reasoning systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21775
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DualSchool: How Reliable are LLMs for Optimization Education?
Klamkin, Michael
Deza, Arnaud
Cheng, Sikai
Zhao, Haoruo
Van Hentenryck, Pascal
Machine Learning
Artificial Intelligence
Optimization and Control
Consider the following task taught in introductory optimization courses which addresses challenges articulated by the community at the intersection of (generative) AI and OR: generate the dual of a linear program. LLMs, being trained at web-scale, have the conversion process and many instances of Primal to Dual Conversion (P2DC) at their disposal. Students may thus reasonably expect that LLMs would perform well on the P2DC task. To assess this expectation, this paper introduces DualSchool, a comprehensive framework for generating and verifying P2DC instances. The verification procedure of DualSchool uses the Canonical Graph Edit Distance, going well beyond existing evaluation methods for optimization models, which exhibit many false positives and negatives when applied to P2DC. Experiments performed by DualSchool reveal interesting findings. Although LLMs can recite the conversion procedure accurately, state-of-the-art open LLMs fail to consistently produce correct duals. This finding holds even for the smallest two-variable instances and for derivative tasks, such as correctness, verification, and error classification. The paper also discusses the implications for educators, students, and the development of large reasoning systems.
title DualSchool: How Reliable are LLMs for Optimization Education?
topic Machine Learning
Artificial Intelligence
Optimization and Control
url https://arxiv.org/abs/2505.21775