Newton's Lantern: A Reinforcement Learning Framework for Finetuning AC Power Flow Warm Start Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bose, Shourya, Hilmarsson, Helgi, Suri, Dhruv
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914555443019776
author Bose, Shourya
Hilmarsson, Helgi
Suri, Dhruv
author_facet Bose, Shourya
Hilmarsson, Helgi
Suri, Dhruv
contents Neural warm starts can sharply reduce the number of Newton-Raphson iterations required to solve the AC power flow problem, but existing supervised approaches generalize poorly on heavily loaded instances near voltage collapse. We prove a lower bound on the Newton-Raphson iteration count that depends on the direction of the warm start error rather than on its magnitude, and show as a corollary that the bound becomes vacuous as the smallest singular value of the power-flow Jacobian shrinks, identifying the failure mode of supervised regression near the saddle-node bifurcation. Motivated by this analysis, we introduce Newton's Lantern, a finetuning pipeline that combines group relative policy optimization with a learned reward model trained on perturbations of the base model's predictions, using the iteration count itself as the supervisory signal. Across IEEE 118-bus, GOC 500-bus, and GOC 2000-bus benchmarks, Newton's Lantern is the only method that converges on every test snapshot while attaining the smallest mean iteration count.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11102
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Newton's Lantern: A Reinforcement Learning Framework for Finetuning AC Power Flow Warm Start Models
Bose, Shourya
Hilmarsson, Helgi
Suri, Dhruv
Machine Learning
Artificial Intelligence
Systems and Control
Neural warm starts can sharply reduce the number of Newton-Raphson iterations required to solve the AC power flow problem, but existing supervised approaches generalize poorly on heavily loaded instances near voltage collapse. We prove a lower bound on the Newton-Raphson iteration count that depends on the direction of the warm start error rather than on its magnitude, and show as a corollary that the bound becomes vacuous as the smallest singular value of the power-flow Jacobian shrinks, identifying the failure mode of supervised regression near the saddle-node bifurcation. Motivated by this analysis, we introduce Newton's Lantern, a finetuning pipeline that combines group relative policy optimization with a learned reward model trained on perturbations of the base model's predictions, using the iteration count itself as the supervisory signal. Across IEEE 118-bus, GOC 500-bus, and GOC 2000-bus benchmarks, Newton's Lantern is the only method that converges on every test snapshot while attaining the smallest mean iteration count.
title Newton's Lantern: A Reinforcement Learning Framework for Finetuning AC Power Flow Warm Start Models
topic Machine Learning
Artificial Intelligence
Systems and Control
url https://arxiv.org/abs/2605.11102