Accelerating Private Large Transformers Inference through Fine-grained Collaborative Computation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yuntian, Tang, Zhanyong, Lu, Tianpei, Zhang, Bingsheng, Shi, Zhiying, Wang, Zheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911035650211840
author Chen, Yuntian
Tang, Zhanyong
Lu, Tianpei
Zhang, Bingsheng
Shi, Zhiying
Wang, Zheng
author_facet Chen, Yuntian
Tang, Zhanyong
Lu, Tianpei
Zhang, Bingsheng
Shi, Zhiying
Wang, Zheng
contents Homomorphic encryption (HE) and secret sharing (SS) enable computations on encrypted data, providing significant privacy benefits for large transformer-based models (TBM) in sensitive sectors like medicine and finance. However, private TBM inference incurs significant costs due to the coarse-grained application of HE and SS. We present FASTLMPI, a new approach to accelerate private TBM inference through fine-grained computation optimization. Specifically, through the fine-grained co-design of homomorphic encryption and secret sharing, FASTLMPI achieves efficient protocols for matrix multiplication, SoftMax, LayerNorm, and GeLU. In addition, FASTLMPI introduces a precise segmented approximation technique for differentiable non-linear, improving its fitting accuracy while maintaining a low polynomial degree. Compared to solution BOLT (S&P'24), FASTLMPI shows a remarkable 54% to 64% decrease in runtime and an impressive 72.2% reduction in communication costs.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16537
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Accelerating Private Large Transformers Inference through Fine-grained Collaborative Computation
Chen, Yuntian
Tang, Zhanyong
Lu, Tianpei
Zhang, Bingsheng
Shi, Zhiying
Wang, Zheng
Cryptography and Security
Homomorphic encryption (HE) and secret sharing (SS) enable computations on encrypted data, providing significant privacy benefits for large transformer-based models (TBM) in sensitive sectors like medicine and finance. However, private TBM inference incurs significant costs due to the coarse-grained application of HE and SS. We present FASTLMPI, a new approach to accelerate private TBM inference through fine-grained computation optimization. Specifically, through the fine-grained co-design of homomorphic encryption and secret sharing, FASTLMPI achieves efficient protocols for matrix multiplication, SoftMax, LayerNorm, and GeLU. In addition, FASTLMPI introduces a precise segmented approximation technique for differentiable non-linear, improving its fitting accuracy while maintaining a low polynomial degree. Compared to solution BOLT (S&P'24), FASTLMPI shows a remarkable 54% to 64% decrease in runtime and an impressive 72.2% reduction in communication costs.
title Accelerating Private Large Transformers Inference through Fine-grained Collaborative Computation
topic Cryptography and Security
url https://arxiv.org/abs/2412.16537