From Classical Data to Quantum Advantage -- Quantum Policy Evaluation on Quantum Hardware

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hein, Daniel, Wiedemann, Simon, Baumann, Markus, Felbinger, Patrik, Klein, Justin, Schieder, Maximilian, Stein, Jonas, Schuman, Daniëlle, Cope, Thomas, Udluft, Steffen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915486464212992
author Hein, Daniel
Wiedemann, Simon
Baumann, Markus
Felbinger, Patrik
Klein, Justin
Schieder, Maximilian
Stein, Jonas
Schuman, Daniëlle
Cope, Thomas
Udluft, Steffen
author_facet Hein, Daniel
Wiedemann, Simon
Baumann, Markus
Felbinger, Patrik
Klein, Justin
Schieder, Maximilian
Stein, Jonas
Schuman, Daniëlle
Cope, Thomas
Udluft, Steffen
contents Quantum policy evaluation (QPE) is a reinforcement learning (RL) algorithm which is quadratically more efficient than an analogous classical Monte Carlo estimation. It makes use of a direct quantum mechanical realization of a finite Markov decision process, in which the agent and the environment are modeled by unitary operators and exchange states, actions, and rewards in superposition. Previously, the quantum environment has been implemented and parametrized manually for an illustrative benchmark using a quantum simulator. In this paper, we demonstrate how these environment parameters can be learned from a batch of classical observational data through quantum machine learning (QML) on quantum hardware. The learned quantum environment is then applied in QPE to also compute policy evaluations on quantum hardware. Our experiments reveal that, despite challenges such as noise and short coherence times, the integration of QML and QPE shows promising potential for achieving quantum advantage in RL.
format Preprint
id arxiv_https___arxiv_org_abs_2509_07614
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Classical Data to Quantum Advantage -- Quantum Policy Evaluation on Quantum Hardware
Hein, Daniel
Wiedemann, Simon
Baumann, Markus
Felbinger, Patrik
Klein, Justin
Schieder, Maximilian
Stein, Jonas
Schuman, Daniëlle
Cope, Thomas
Udluft, Steffen
Quantum Physics
Artificial Intelligence
Quantum policy evaluation (QPE) is a reinforcement learning (RL) algorithm which is quadratically more efficient than an analogous classical Monte Carlo estimation. It makes use of a direct quantum mechanical realization of a finite Markov decision process, in which the agent and the environment are modeled by unitary operators and exchange states, actions, and rewards in superposition. Previously, the quantum environment has been implemented and parametrized manually for an illustrative benchmark using a quantum simulator. In this paper, we demonstrate how these environment parameters can be learned from a batch of classical observational data through quantum machine learning (QML) on quantum hardware. The learned quantum environment is then applied in QPE to also compute policy evaluations on quantum hardware. Our experiments reveal that, despite challenges such as noise and short coherence times, the integration of QML and QPE shows promising potential for achieving quantum advantage in RL.
title From Classical Data to Quantum Advantage -- Quantum Policy Evaluation on Quantum Hardware
topic Quantum Physics
Artificial Intelligence
url https://arxiv.org/abs/2509.07614