Distilling Human-Aligned Privacy Sensitivity Assessment from Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Loiseau, Gabriel, Sileo, Damien, Riquet, Damien, Meyer, Maxime, Tommasi, Marc
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915901387833344
author Loiseau, Gabriel
Sileo, Damien
Riquet, Damien
Meyer, Maxime
Tommasi, Marc
author_facet Loiseau, Gabriel
Sileo, Damien
Riquet, Damien
Meyer, Maxime
Tommasi, Marc
contents Accurate privacy evaluation of textual data remains a critical challenge in privacy-preserving natural language processing. Recent work has shown that large language models (LLMs) can serve as reliable privacy evaluators, achieving strong agreement with human judgments; however, their computational cost and impracticality for processing sensitive data at scale limit real-world deployment. We address this gap by distilling the privacy assessment capabilities of Mistral Large 3 (675B) into lightweight encoder models with as few as 150M parameters. Leveraging a large-scale dataset of privacy-annotated texts spanning 10 diverse domains, we train efficient classifiers that preserve strong agreement with human annotations while dramatically reducing computational requirements. We validate our approach on human-annotated test data and demonstrate its practical utility as an evaluation metric for de-identification systems.
format Preprint
id arxiv_https___arxiv_org_abs_2603_29497
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Distilling Human-Aligned Privacy Sensitivity Assessment from Large Language Models
Loiseau, Gabriel
Sileo, Damien
Riquet, Damien
Meyer, Maxime
Tommasi, Marc
Computation and Language
Accurate privacy evaluation of textual data remains a critical challenge in privacy-preserving natural language processing. Recent work has shown that large language models (LLMs) can serve as reliable privacy evaluators, achieving strong agreement with human judgments; however, their computational cost and impracticality for processing sensitive data at scale limit real-world deployment. We address this gap by distilling the privacy assessment capabilities of Mistral Large 3 (675B) into lightweight encoder models with as few as 150M parameters. Leveraging a large-scale dataset of privacy-annotated texts spanning 10 diverse domains, we train efficient classifiers that preserve strong agreement with human annotations while dramatically reducing computational requirements. We validate our approach on human-annotated test data and demonstrate its practical utility as an evaluation metric for de-identification systems.
title Distilling Human-Aligned Privacy Sensitivity Assessment from Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2603.29497