Preempting Text Sanitization Utility in Resource-Constrained Privacy-Preserving LLM Interactions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Carpentier, Robin, Zhao, Benjamin Zi Hao, Asghar, Hassan Jameel, Kaafar, Dali
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909647883993088
author Carpentier, Robin
Zhao, Benjamin Zi Hao
Asghar, Hassan Jameel
Kaafar, Dali
author_facet Carpentier, Robin
Zhao, Benjamin Zi Hao
Asghar, Hassan Jameel
Kaafar, Dali
contents Interactions with online Large Language Models raise privacy issues where providers can gather sensitive information about users and their companies from the prompts. While textual prompts can be sanitized using Differential Privacy, we show that it is difficult to anticipate the performance of an LLM on such sanitized prompt. Poor performance has clear monetary consequences for LLM services charging on a pay-per-use model as well as great amount of computing resources wasted. To this end, we propose a middleware architecture leveraging a Small Language Model to predict the utility of a given sanitized prompt before it is sent to the LLM. We experimented on a summarization task and a translation task to show that our architecture helps prevent such resource waste for up to 20% of the prompts. During our study, we also reproduced experiments from one of the most cited paper on text sanitization using DP and show that a potential performance-driven implementation choice dramatically changes the output while not being explicitly acknowledged in the paper.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11521
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Preempting Text Sanitization Utility in Resource-Constrained Privacy-Preserving LLM Interactions
Carpentier, Robin
Zhao, Benjamin Zi Hao
Asghar, Hassan Jameel
Kaafar, Dali
Cryptography and Security
Machine Learning
Interactions with online Large Language Models raise privacy issues where providers can gather sensitive information about users and their companies from the prompts. While textual prompts can be sanitized using Differential Privacy, we show that it is difficult to anticipate the performance of an LLM on such sanitized prompt. Poor performance has clear monetary consequences for LLM services charging on a pay-per-use model as well as great amount of computing resources wasted. To this end, we propose a middleware architecture leveraging a Small Language Model to predict the utility of a given sanitized prompt before it is sent to the LLM. We experimented on a summarization task and a translation task to show that our architecture helps prevent such resource waste for up to 20% of the prompts. During our study, we also reproduced experiments from one of the most cited paper on text sanitization using DP and show that a potential performance-driven implementation choice dramatically changes the output while not being explicitly acknowledged in the paper.
title Preempting Text Sanitization Utility in Resource-Constrained Privacy-Preserving LLM Interactions
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2411.11521