VQC-Based Reinforcement Learning with Data Re-uploading: Performance and Trainability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Coelho, Rodrigo, Sequeira, André, Santos, Luís Paulo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909384837169152
author Coelho, Rodrigo
Sequeira, André
Santos, Luís Paulo
author_facet Coelho, Rodrigo
Sequeira, André
Santos, Luís Paulo
contents Reinforcement Learning (RL) consists of designing agents that make intelligent decisions without human supervision. When used alongside function approximators such as Neural Networks (NNs), RL is capable of solving extremely complex problems. Deep Q-Learning, a RL algorithm that uses Deep NNs, achieved super-human performance in some specific tasks. Nonetheless, it is also possible to use Variational Quantum Circuits (VQCs) as function approximators in RL algorithms. This work empirically studies the performance and trainability of such VQC-based Deep Q-Learning models in classic control benchmark environments. More specifically, we research how data re-uploading affects both these metrics. We show that the magnitude and the variance of the gradients of these models remain substantial throughout training due to the moving targets of Deep Q-Learning. Moreover, we empirically show that increasing the number of qubits does not lead to an exponential vanishing behavior of the magnitude and variance of the gradients for a PQC approximating a 2-design, unlike what was expected due to the Barren Plateau Phenomenon. This hints at the possibility of VQCs being specially adequate for being used as function approximators in such a context.
format Preprint
id arxiv_https___arxiv_org_abs_2401_11555
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VQC-Based Reinforcement Learning with Data Re-uploading: Performance and Trainability
Coelho, Rodrigo
Sequeira, André
Santos, Luís Paulo
Quantum Physics
Machine Learning
Reinforcement Learning (RL) consists of designing agents that make intelligent decisions without human supervision. When used alongside function approximators such as Neural Networks (NNs), RL is capable of solving extremely complex problems. Deep Q-Learning, a RL algorithm that uses Deep NNs, achieved super-human performance in some specific tasks. Nonetheless, it is also possible to use Variational Quantum Circuits (VQCs) as function approximators in RL algorithms. This work empirically studies the performance and trainability of such VQC-based Deep Q-Learning models in classic control benchmark environments. More specifically, we research how data re-uploading affects both these metrics. We show that the magnitude and the variance of the gradients of these models remain substantial throughout training due to the moving targets of Deep Q-Learning. Moreover, we empirically show that increasing the number of qubits does not lead to an exponential vanishing behavior of the magnitude and variance of the gradients for a PQC approximating a 2-design, unlike what was expected due to the Barren Plateau Phenomenon. This hints at the possibility of VQCs being specially adequate for being used as function approximators in such a context.
title VQC-Based Reinforcement Learning with Data Re-uploading: Performance and Trainability
topic Quantum Physics
Machine Learning
url https://arxiv.org/abs/2401.11555