Quantamination: Dynamic Quantization Leaks Your Data Across the Batch

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Foerster, Hanna, Shumailov, Ilia, Zhang, Cheng, Zhao, Yiren, Hayes, Jamie, Mullins, Robert
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911632514351104
author Foerster, Hanna
Shumailov, Ilia
Zhang, Cheng
Zhao, Yiren
Hayes, Jamie
Mullins, Robert
author_facet Foerster, Hanna
Shumailov, Ilia
Zhang, Cheng
Zhao, Yiren
Hayes, Jamie
Mullins, Robert
contents Dynamic quantization emerged as a practical approach to increase the utilization and efficiency of the machine learning serving flow. Unlike static quantization, which applies quantization offline, dynamic quantization operates on tensors at run-time, adapting its parameters to the actual input data. Today's mainstream machine learning frameworks, including ML compilers and inference engines, frequently recommend dynamic quantization as an initial step for optimizing model serving. This is because dynamic quantization can significantly reduce memory usage and computational load, leading to faster token generation and improved model serving efficiency without substantial loss in model accuracy. In this paper, we reveal a critical vulnerability in dynamic quantization: an adversary can exploit such quantization strategy to steal sensitive user data placed in the same batch as the adversary's input. Our analysis demonstrates that dynamic quantization, when improperly implemented or configured, can create side channels that expose information about other inputs within the same batch. We call this phenomenon Quantamination, describing contamination from quantization. Specifically, we show that at least 4 of the most popular ML frameworks in use today either default to or can use configurations that leak data across the batch boundary. This data leakage, in theory, allows attackers to partially or even fully recover other users' batched input data, representing a serious privacy risk for existing ML serving frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_26505
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
Foerster, Hanna
Shumailov, Ilia
Zhang, Cheng
Zhao, Yiren
Hayes, Jamie
Mullins, Robert
Cryptography and Security
Machine Learning
Dynamic quantization emerged as a practical approach to increase the utilization and efficiency of the machine learning serving flow. Unlike static quantization, which applies quantization offline, dynamic quantization operates on tensors at run-time, adapting its parameters to the actual input data. Today's mainstream machine learning frameworks, including ML compilers and inference engines, frequently recommend dynamic quantization as an initial step for optimizing model serving. This is because dynamic quantization can significantly reduce memory usage and computational load, leading to faster token generation and improved model serving efficiency without substantial loss in model accuracy. In this paper, we reveal a critical vulnerability in dynamic quantization: an adversary can exploit such quantization strategy to steal sensitive user data placed in the same batch as the adversary's input. Our analysis demonstrates that dynamic quantization, when improperly implemented or configured, can create side channels that expose information about other inputs within the same batch. We call this phenomenon Quantamination, describing contamination from quantization. Specifically, we show that at least 4 of the most popular ML frameworks in use today either default to or can use configurations that leak data across the batch boundary. This data leakage, in theory, allows attackers to partially or even fully recover other users' batched input data, representing a serious privacy risk for existing ML serving frameworks.
title Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2604.26505