Nimbus: Secure and Efficient Two-Party Inference for Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhengyi, Yang, Kang, Tan, Jin, Lu, Wen-jie, Wu, Haoqi, Wang, Xiao, Yu, Yu, Zhao, Derun, Zheng, Yancheng, Guo, Minyi, Leng, Jingwen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916494017822720
author Li, Zhengyi
Yang, Kang
Tan, Jin
Lu, Wen-jie
Wu, Haoqi
Wang, Xiao
Yu, Yu
Zhao, Derun
Zheng, Yancheng
Guo, Minyi
Leng, Jingwen
author_facet Li, Zhengyi
Yang, Kang
Tan, Jin
Lu, Wen-jie
Wu, Haoqi
Wang, Xiao
Yu, Yu
Zhao, Derun
Zheng, Yancheng
Guo, Minyi
Leng, Jingwen
contents Transformer models have gained significant attention due to their power in machine learning tasks. Their extensive deployment has raised concerns about the potential leakage of sensitive information during inference. However, when being applied to Transformers, existing approaches based on secure two-party computation (2PC) bring about efficiency limitations in two folds: (1) resource-intensive matrix multiplications in linear layers, and (2) complex non-linear activation functions like $\mathsf{GELU}$ and $\mathsf{Softmax}$. This work presents a new two-party inference framework $\mathsf{Nimbus}$ for Transformer models. For the linear layer, we propose a new 2PC paradigm along with an encoding approach to securely compute matrix multiplications based on an outer-product insight, which achieves $2.9\times \sim 12.5\times$ performance improvements compared to the state-of-the-art (SOTA) protocol. For the non-linear layer, through a new observation of utilizing the input distribution, we propose an approach of low-degree polynomial approximation for $\mathsf{GELU}$ and $\mathsf{Softmax}$, which improves the performance of the SOTA polynomial approximation by $2.9\times \sim 4.0\times$, where the average accuracy loss of our approach is 0.08\% compared to the non-2PC inference without privacy. Compared with the SOTA two-party inference, $\mathsf{Nimbus}$ improves the end-to-end performance of \bert{} inference by $2.7\times \sim 4.7\times$ across different network settings.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15707
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Nimbus: Secure and Efficient Two-Party Inference for Transformers
Li, Zhengyi
Yang, Kang
Tan, Jin
Lu, Wen-jie
Wu, Haoqi
Wang, Xiao
Yu, Yu
Zhao, Derun
Zheng, Yancheng
Guo, Minyi
Leng, Jingwen
Cryptography and Security
Artificial Intelligence
Transformer models have gained significant attention due to their power in machine learning tasks. Their extensive deployment has raised concerns about the potential leakage of sensitive information during inference. However, when being applied to Transformers, existing approaches based on secure two-party computation (2PC) bring about efficiency limitations in two folds: (1) resource-intensive matrix multiplications in linear layers, and (2) complex non-linear activation functions like $\mathsf{GELU}$ and $\mathsf{Softmax}$. This work presents a new two-party inference framework $\mathsf{Nimbus}$ for Transformer models. For the linear layer, we propose a new 2PC paradigm along with an encoding approach to securely compute matrix multiplications based on an outer-product insight, which achieves $2.9\times \sim 12.5\times$ performance improvements compared to the state-of-the-art (SOTA) protocol. For the non-linear layer, through a new observation of utilizing the input distribution, we propose an approach of low-degree polynomial approximation for $\mathsf{GELU}$ and $\mathsf{Softmax}$, which improves the performance of the SOTA polynomial approximation by $2.9\times \sim 4.0\times$, where the average accuracy loss of our approach is 0.08\% compared to the non-2PC inference without privacy. Compared with the SOTA two-party inference, $\mathsf{Nimbus}$ improves the end-to-end performance of \bert{} inference by $2.7\times \sim 4.7\times$ across different network settings.
title Nimbus: Secure and Efficient Two-Party Inference for Transformers
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2411.15707