Quantum Large Language Model Fine-Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Sang Hyub, Mei, Jonathan, Girotto, Claudio, Yamada, Masako, Roetteler, Martin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910909319872512
author Kim, Sang Hyub
Mei, Jonathan
Girotto, Claudio
Yamada, Masako
Roetteler, Martin
author_facet Kim, Sang Hyub
Mei, Jonathan
Girotto, Claudio
Yamada, Masako
Roetteler, Martin
contents We introduce a hybrid quantum-classical deep learning architecture for large language model fine-tuning. The classical portion of the architecture is a sentence transformer that is powerful enough to display significant accuracy for complex tasks such as sentiment prediction. The quantum portion of the architecture consists of parameterized quantum circuits that utilize long-range connections between qubits. We analyze the performance of the hybrid models for various settings of hyperparameters, including the number of qubits, the depth of the quantum circuits, learning rate, number of re-uploading steps, etc. Based on a screening study of main effects, we show an overall improvement in prediction accuracy over a comparable classical baseline, with a trend of increasing accuracy with number of qubits. We observe up to $3.14\%$ improvements in accuracy over classical architectures of comparable model size, within the set of hyperparameters probed in this study. We demonstrate the contribution of each module in our architecture through ablation studies. Our studies are based on finite shot-counts and include simulations based on noisy quantum gates.
format Preprint
id arxiv_https___arxiv_org_abs_2504_08732
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Quantum Large Language Model Fine-Tuning
Kim, Sang Hyub
Mei, Jonathan
Girotto, Claudio
Yamada, Masako
Roetteler, Martin
Quantum Physics
Emerging Technologies
We introduce a hybrid quantum-classical deep learning architecture for large language model fine-tuning. The classical portion of the architecture is a sentence transformer that is powerful enough to display significant accuracy for complex tasks such as sentiment prediction. The quantum portion of the architecture consists of parameterized quantum circuits that utilize long-range connections between qubits. We analyze the performance of the hybrid models for various settings of hyperparameters, including the number of qubits, the depth of the quantum circuits, learning rate, number of re-uploading steps, etc. Based on a screening study of main effects, we show an overall improvement in prediction accuracy over a comparable classical baseline, with a trend of increasing accuracy with number of qubits. We observe up to $3.14\%$ improvements in accuracy over classical architectures of comparable model size, within the set of hyperparameters probed in this study. We demonstrate the contribution of each module in our architecture through ablation studies. Our studies are based on finite shot-counts and include simulations based on noisy quantum gates.
title Quantum Large Language Model Fine-Tuning
topic Quantum Physics
Emerging Technologies
url https://arxiv.org/abs/2504.08732