Q-Policy: Quantum-Enhanced Policy Evaluation for Scalable Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cherukuri, Kalyan, Lala, Aarav, Yardi, Yash
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913882429194240
author Cherukuri, Kalyan
Lala, Aarav
Yardi, Yash
author_facet Cherukuri, Kalyan
Lala, Aarav
Yardi, Yash
contents We propose Q-Policy, a hybrid quantum-classical reinforcement learning (RL) framework that mathematically accelerates policy evaluation and optimization by exploiting quantum computing primitives. Q-Policy encodes value functions in quantum superposition, enabling simultaneous evaluation of multiple state-action pairs via amplitude encoding and quantum parallelism. We introduce a quantum-enhanced policy iteration algorithm with provable polynomial reductions in sample complexity for the evaluation step, under standard assumptions. To demonstrate the technical feasibility and theoretical soundness of our approach, we validate Q-Policy on classical emulations of small discrete control tasks. Due to current hardware and simulation limitations, our experiments focus on showcasing proof-of-concept behavior rather than large-scale empirical evaluation. Our results support the potential of Q-Policy as a theoretical foundation for scalable RL on future quantum devices, addressing RL scalability challenges beyond classical approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11862
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Q-Policy: Quantum-Enhanced Policy Evaluation for Scalable Reinforcement Learning
Cherukuri, Kalyan
Lala, Aarav
Yardi, Yash
Machine Learning
Artificial Intelligence
Quantum Physics
We propose Q-Policy, a hybrid quantum-classical reinforcement learning (RL) framework that mathematically accelerates policy evaluation and optimization by exploiting quantum computing primitives. Q-Policy encodes value functions in quantum superposition, enabling simultaneous evaluation of multiple state-action pairs via amplitude encoding and quantum parallelism. We introduce a quantum-enhanced policy iteration algorithm with provable polynomial reductions in sample complexity for the evaluation step, under standard assumptions. To demonstrate the technical feasibility and theoretical soundness of our approach, we validate Q-Policy on classical emulations of small discrete control tasks. Due to current hardware and simulation limitations, our experiments focus on showcasing proof-of-concept behavior rather than large-scale empirical evaluation. Our results support the potential of Q-Policy as a theoretical foundation for scalable RL on future quantum devices, addressing RL scalability challenges beyond classical approaches.
title Q-Policy: Quantum-Enhanced Policy Evaluation for Scalable Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Quantum Physics
url https://arxiv.org/abs/2505.11862