Saved in:
Bibliographic Details
Main Authors: Angioli, Marco, Johansson, Kevin, Rosato, Antonello, Loutfi, Amy, Kleyko, Denis
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.13346
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918499608166400
author Angioli, Marco
Johansson, Kevin
Rosato, Antonello
Loutfi, Amy
Kleyko, Denis
author_facet Angioli, Marco
Johansson, Kevin
Rosato, Antonello
Loutfi, Amy
Kleyko, Denis
contents Contextual bandits (CB) are online sequential decision-making problems under partial feedback that underpin many adaptive services. There is a growing demand to deploy CB agents directly on-device, under strict constraints on memory, compute, and energy. However, standard linear CB algorithms are often impractical for resource-constrained devices with their unfavorable scaling in computational and memory costs. Recently, HD-CB, a CB approach based on hyperdimensional computing principles, has been proposed to model and solve CB problems by moving into high-dimensional spaces. HD-CB offers faster convergence, favorable scalability, and improves memory efficiency compared to linear CB algorithms. However, its learning rule is accumulation-based: the values of action vectors grow over time, requiring high precision. While periodic binarization can prevent overflow in low-precision components, it may discard important information about magnitudes and degrade decision quality. This paper introduces probabilistic HD-CB, a low-precision variant that replaces deterministic accumulation with a probabilistic update rule. At each step, only a random subset of vector components is updated, with a time-decaying update probability, and component values are constrained to a predefined range [-k,+k]. This approach enables low-precision components, prevents overflow without periodic binarization, and reduces the expected update cost in proportion to the fraction of updated components. Off-policy evaluation on standardized synthetic CB benchmarks using the Open Bandit Pipeline shows that probabilistic HD-CB consistently outperforms binarized HD-CB at equal precision, while approaching the performance of HD-CB with as few as 3 bits per component.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13346
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Contextual Bandits for Resource-Constrained Devices using Probabilistic Learning
Angioli, Marco
Johansson, Kevin
Rosato, Antonello
Loutfi, Amy
Kleyko, Denis
Machine Learning
Contextual bandits (CB) are online sequential decision-making problems under partial feedback that underpin many adaptive services. There is a growing demand to deploy CB agents directly on-device, under strict constraints on memory, compute, and energy. However, standard linear CB algorithms are often impractical for resource-constrained devices with their unfavorable scaling in computational and memory costs. Recently, HD-CB, a CB approach based on hyperdimensional computing principles, has been proposed to model and solve CB problems by moving into high-dimensional spaces. HD-CB offers faster convergence, favorable scalability, and improves memory efficiency compared to linear CB algorithms. However, its learning rule is accumulation-based: the values of action vectors grow over time, requiring high precision. While periodic binarization can prevent overflow in low-precision components, it may discard important information about magnitudes and degrade decision quality. This paper introduces probabilistic HD-CB, a low-precision variant that replaces deterministic accumulation with a probabilistic update rule. At each step, only a random subset of vector components is updated, with a time-decaying update probability, and component values are constrained to a predefined range [-k,+k]. This approach enables low-precision components, prevents overflow without periodic binarization, and reduces the expected update cost in proportion to the fraction of updated components. Off-policy evaluation on standardized synthetic CB benchmarks using the Open Bandit Pipeline shows that probabilistic HD-CB consistently outperforms binarized HD-CB at equal precision, while approaching the performance of HD-CB with as few as 3 bits per component.
title Contextual Bandits for Resource-Constrained Devices using Probabilistic Learning
topic Machine Learning
url https://arxiv.org/abs/2605.13346