ProOPF: Benchmarking and Improving LLMs for Professional-Grade Power Systems Optimization Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Chao, Guo, Zihan, Wan, Xu, Yang, Zhenghao, Zhang, Yifan, Huang, Wengi, Song, Jie, Zhang, Zongyan, Sun, Mingyang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913153635319808
author Shen, Chao
Guo, Zihan
Wan, Xu
Yang, Zhenghao
Zhang, Yifan
Huang, Wengi
Song, Jie
Zhang, Zongyan
Sun, Mingyang
author_facet Shen, Chao
Guo, Zihan
Wan, Xu
Yang, Zhenghao
Zhang, Yifan
Huang, Wengi
Song, Jie
Zhang, Zongyan
Sun, Mingyang
contents Growing renewable penetration introduces substantial uncertainty into power system operations, necessitating frequent adaptation of dispatch objectives and constraints and challenging expertise-intensive, near-real-time modeling workflows. Large Language Models (LLMs) provide a promising avenue for automating this process by translating natural-language (NL) operational requirements into executable optimization models via semantic reasoning and code synthesis. Yet existing LLM datasets and benchmarks for optimization modeling primarily target coarse-grained cross-domain generalization, offering limited, rigorous evaluation in power-system settings, particularly for Optimal Power Flow (OPF). We therefore introduce \textbf{ProOPF-D} and \textbf{ProOPF-B}, a dataset and benchmark for professional-grade OPF modeling: ProOPF-D contains 12K instances pairing NL requests with parameter adjustments and structural extensions to a canonical OPF, together with executable implementations; ProOPF-B provides 121 expert-annotated test cases with ground-truth code, enabling end-to-end evaluation under both concrete and abstract OPF modeling regimes.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03070
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ProOPF: Benchmarking and Improving LLMs for Professional-Grade Power Systems Optimization Modeling
Shen, Chao
Guo, Zihan
Wan, Xu
Yang, Zhenghao
Zhang, Yifan
Huang, Wengi
Song, Jie
Zhang, Zongyan
Sun, Mingyang
Systems and Control
Software Engineering
Growing renewable penetration introduces substantial uncertainty into power system operations, necessitating frequent adaptation of dispatch objectives and constraints and challenging expertise-intensive, near-real-time modeling workflows. Large Language Models (LLMs) provide a promising avenue for automating this process by translating natural-language (NL) operational requirements into executable optimization models via semantic reasoning and code synthesis. Yet existing LLM datasets and benchmarks for optimization modeling primarily target coarse-grained cross-domain generalization, offering limited, rigorous evaluation in power-system settings, particularly for Optimal Power Flow (OPF). We therefore introduce \textbf{ProOPF-D} and \textbf{ProOPF-B}, a dataset and benchmark for professional-grade OPF modeling: ProOPF-D contains 12K instances pairing NL requests with parameter adjustments and structural extensions to a canonical OPF, together with executable implementations; ProOPF-B provides 121 expert-annotated test cases with ground-truth code, enabling end-to-end evaluation under both concrete and abstract OPF modeling regimes.
title ProOPF: Benchmarking and Improving LLMs for Professional-Grade Power Systems Optimization Modeling
topic Systems and Control
Software Engineering
url https://arxiv.org/abs/2602.03070