Quantum-Audit: Evaluating the Reasoning Limits of LLMs on Quantum Computing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Afane, Mohamed, Laufer, Kayla, Wei, Wenqi, Mao, Ying, Farooq, Junaid, Wang, Ying, Chen, Juntao
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908826588938240
author Afane, Mohamed
Laufer, Kayla
Wei, Wenqi
Mao, Ying
Farooq, Junaid
Wang, Ying
Chen, Juntao
author_facet Afane, Mohamed
Laufer, Kayla
Wei, Wenqi
Mao, Ying
Farooq, Junaid
Wang, Ying
Chen, Juntao
contents Language models have become practical tools for quantum computing education and research, from summarizing technical papers to explaining theoretical concepts and answering questions about recent developments in the field. While existing benchmarks evaluate quantum code generation and circuit design, their understanding of quantum computing concepts has not been systematically measured. Quantum-Audit addresses this gap with 2,700 questions covering core quantum computing topics. We evaluate 26 models from leading organizations. Our benchmark comprises 1,000 expert-written questions, 1,000 questions extracted from research papers using LLMs and validated by experts, plus an additional 700 questions including 350 open-ended questions and 350 questions with false premises to test whether models can correct erroneous assumptions. Human participants scored between 23% and 86%, with experts averaging 74%. Top-performing models exceeded the expert average, with Claude Opus 4.5 reaching 84% accuracy, though top models showed an average 12-point accuracy drop on expert-written questions compared to LLM-generated ones. Performance declined further on advanced topics, dropping to 73% on security questions. Additionally, models frequently accepted and reinforced false premises embedded in questions instead of identifying them, with accuracy below 66% on these critical reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10092
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Quantum-Audit: Evaluating the Reasoning Limits of LLMs on Quantum Computing
Afane, Mohamed
Laufer, Kayla
Wei, Wenqi
Mao, Ying
Farooq, Junaid
Wang, Ying
Chen, Juntao
Computation and Language
Language models have become practical tools for quantum computing education and research, from summarizing technical papers to explaining theoretical concepts and answering questions about recent developments in the field. While existing benchmarks evaluate quantum code generation and circuit design, their understanding of quantum computing concepts has not been systematically measured. Quantum-Audit addresses this gap with 2,700 questions covering core quantum computing topics. We evaluate 26 models from leading organizations. Our benchmark comprises 1,000 expert-written questions, 1,000 questions extracted from research papers using LLMs and validated by experts, plus an additional 700 questions including 350 open-ended questions and 350 questions with false premises to test whether models can correct erroneous assumptions. Human participants scored between 23% and 86%, with experts averaging 74%. Top-performing models exceeded the expert average, with Claude Opus 4.5 reaching 84% accuracy, though top models showed an average 12-point accuracy drop on expert-written questions compared to LLM-generated ones. Performance declined further on advanced topics, dropping to 73% on security questions. Additionally, models frequently accepted and reinforced false premises embedded in questions instead of identifying them, with accuracy below 66% on these critical reasoning tasks.
title Quantum-Audit: Evaluating the Reasoning Limits of LLMs on Quantum Computing
topic Computation and Language
url https://arxiv.org/abs/2602.10092