Practical and Private Hybrid ML Inference with Fully Homomorphic Encryption

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Biswas, Sayan, Chartier, Philippe, Dhasade, Akash, Jurien, Tom, Kerriou, David, Kerrmarec, Anne-Marie, Lemou, Mohammed, Tranie, Franklin, de Vos, Martijn, Vujasinovic, Milos
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914015555354624
author Biswas, Sayan
Chartier, Philippe
Dhasade, Akash
Jurien, Tom
Kerriou, David
Kerrmarec, Anne-Marie
Lemou, Mohammed
Tranie, Franklin
de Vos, Martijn
Vujasinovic, Milos
author_facet Biswas, Sayan
Chartier, Philippe
Dhasade, Akash
Jurien, Tom
Kerriou, David
Kerrmarec, Anne-Marie
Lemou, Mohammed
Tranie, Franklin
de Vos, Martijn
Vujasinovic, Milos
contents In contemporary cloud-based services, protecting users' sensitive data and ensuring the confidentiality of the server's model are critical. Fully homomorphic encryption (FHE) enables inference directly on encrypted inputs, but its practicality is hindered by expensive bootstrapping and inefficient approximations of non-linear activations. We introduce Safhire, a hybrid inference framework that executes linear layers under encryption on the server while offloading non-linearities to the client in plaintext. This design eliminates bootstrapping, supports exact activations, and significantly reduces computation. To safeguard model confidentiality despite client access to intermediate outputs, Safhire applies randomized shuffling, which obfuscates intermediate values and makes it practically impossible to reconstruct the model. To further reduce latency, Safhire incorporates advanced optimizations such as fast ciphertext packing and partial extraction. Evaluations on multiple standard models and datasets show that Safhire achieves 1.5X - 10.5X lower inference latency than Orion, a state-of-the-art baseline, with manageable communication overhead and comparable accuracy, thereby establishing the practicality of hybrid FHE inference.
format Preprint
id arxiv_https___arxiv_org_abs_2509_01253
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Practical and Private Hybrid ML Inference with Fully Homomorphic Encryption
Biswas, Sayan
Chartier, Philippe
Dhasade, Akash
Jurien, Tom
Kerriou, David
Kerrmarec, Anne-Marie
Lemou, Mohammed
Tranie, Franklin
de Vos, Martijn
Vujasinovic, Milos
Cryptography and Security
Machine Learning
In contemporary cloud-based services, protecting users' sensitive data and ensuring the confidentiality of the server's model are critical. Fully homomorphic encryption (FHE) enables inference directly on encrypted inputs, but its practicality is hindered by expensive bootstrapping and inefficient approximations of non-linear activations. We introduce Safhire, a hybrid inference framework that executes linear layers under encryption on the server while offloading non-linearities to the client in plaintext. This design eliminates bootstrapping, supports exact activations, and significantly reduces computation. To safeguard model confidentiality despite client access to intermediate outputs, Safhire applies randomized shuffling, which obfuscates intermediate values and makes it practically impossible to reconstruct the model. To further reduce latency, Safhire incorporates advanced optimizations such as fast ciphertext packing and partial extraction. Evaluations on multiple standard models and datasets show that Safhire achieves 1.5X - 10.5X lower inference latency than Orion, a state-of-the-art baseline, with manageable communication overhead and comparable accuracy, thereby establishing the practicality of hybrid FHE inference.
title Practical and Private Hybrid ML Inference with Fully Homomorphic Encryption
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2509.01253