Can LLMs Solve Science or Just Write Code? Evaluating Quantum Solver Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baresi, Luciano, Bianculli, Domenico, Ernzer, Maryse, Lestingi, Livia, Pastore, Fabrizio, Shin, Seung Yeob
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917480523366400
author Baresi, Luciano
Bianculli, Domenico
Ernzer, Maryse
Lestingi, Livia
Pastore, Fabrizio
Shin, Seung Yeob
author_facet Baresi, Luciano
Bianculli, Domenico
Ernzer, Maryse
Lestingi, Livia
Pastore, Fabrizio
Shin, Seung Yeob
contents Large Language Models (LLMs) show strong capabilities in code generation, motivating their use in automated quantum solver development. However, in quantum computing, successful execution of generated code is not sufficient: correctness depends on numerically accurate results, which are sensitive to non-trivial mappings, hybrid quantum-classical workflows, and algorithm-specific approximations. This work introduces Q-SAGE, an iterative methodology to evaluate LLMs' capability in generating quantum solvers for scientific problems. The methodology adopts an iterative approach by executing the script generated by the LLM, comparing the result with the result of a classical solver, and refining the script until the two results match within a tolerance threshold. We empirically evaluated the methodology with five families of scientific problems of different complexities and five LLMs, both open source and proprietary. The results show that iterative refinement substantially improves success rates, but introduces a significant computational overhead. Moreover, as model capability increases, failure modes shift from execution errors to numerical inaccuracies, highlighting the current limitations of LLM-based quantum software.
format Preprint
id arxiv_https___arxiv_org_abs_2605_07525
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Can LLMs Solve Science or Just Write Code? Evaluating Quantum Solver Generation
Baresi, Luciano
Bianculli, Domenico
Ernzer, Maryse
Lestingi, Livia
Pastore, Fabrizio
Shin, Seung Yeob
Software Engineering
Large Language Models (LLMs) show strong capabilities in code generation, motivating their use in automated quantum solver development. However, in quantum computing, successful execution of generated code is not sufficient: correctness depends on numerically accurate results, which are sensitive to non-trivial mappings, hybrid quantum-classical workflows, and algorithm-specific approximations. This work introduces Q-SAGE, an iterative methodology to evaluate LLMs' capability in generating quantum solvers for scientific problems. The methodology adopts an iterative approach by executing the script generated by the LLM, comparing the result with the result of a classical solver, and refining the script until the two results match within a tolerance threshold. We empirically evaluated the methodology with five families of scientific problems of different complexities and five LLMs, both open source and proprietary. The results show that iterative refinement substantially improves success rates, but introduces a significant computational overhead. Moreover, as model capability increases, failure modes shift from execution errors to numerical inaccuracies, highlighting the current limitations of LLM-based quantum software.
title Can LLMs Solve Science or Just Write Code? Evaluating Quantum Solver Generation
topic Software Engineering
url https://arxiv.org/abs/2605.07525