Comet: A Communication-efficient and Performant Approximation for Private Transformer Inference

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Xiangrui, Zhang, Qiao, Ning, Rui, Xin, Chunsheng, Wu, Hongyi
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913493125431296
author Xu, Xiangrui
Zhang, Qiao
Ning, Rui
Xin, Chunsheng
Wu, Hongyi
author_facet Xu, Xiangrui
Zhang, Qiao
Ning, Rui
Xin, Chunsheng
Wu, Hongyi
contents The prevalent use of Transformer-like models, exemplified by ChatGPT in modern language processing applications, underscores the critical need for enabling private inference essential for many cloud-based services reliant on such models. However, current privacy-preserving frameworks impose significant communication burden, especially for non-linear computation in Transformer model. In this paper, we introduce a novel plug-in method Comet to effectively reduce the communication cost without compromising the inference performance. We second introduce an efficient approximation method to eliminate the heavy communication in finding good initial approximation. We evaluate our Comet on Bert and RoBERTa models with GLUE benchmark datasets, showing up to 3.9$\times$ less communication and 3.5$\times$ speedups while keep competitive model performance compared to the prior art.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17485
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Comet: A Communication-efficient and Performant Approximation for Private Transformer Inference
Xu, Xiangrui
Zhang, Qiao
Ning, Rui
Xin, Chunsheng
Wu, Hongyi
Machine Learning
Artificial Intelligence
Cryptography and Security
The prevalent use of Transformer-like models, exemplified by ChatGPT in modern language processing applications, underscores the critical need for enabling private inference essential for many cloud-based services reliant on such models. However, current privacy-preserving frameworks impose significant communication burden, especially for non-linear computation in Transformer model. In this paper, we introduce a novel plug-in method Comet to effectively reduce the communication cost without compromising the inference performance. We second introduce an efficient approximation method to eliminate the heavy communication in finding good initial approximation. We evaluate our Comet on Bert and RoBERTa models with GLUE benchmark datasets, showing up to 3.9$\times$ less communication and 3.5$\times$ speedups while keep competitive model performance compared to the prior art.
title Comet: A Communication-efficient and Performant Approximation for Private Transformer Inference
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2405.17485