Distilling Human-Aligned Privacy Sensitivity Assessment from Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915901387833344 |
|---|---|
| author | Loiseau, Gabriel Sileo, Damien Riquet, Damien Meyer, Maxime Tommasi, Marc |
| author_facet | Loiseau, Gabriel Sileo, Damien Riquet, Damien Meyer, Maxime Tommasi, Marc |
| contents | Accurate privacy evaluation of textual data remains a critical challenge in privacy-preserving natural language processing. Recent work has shown that large language models (LLMs) can serve as reliable privacy evaluators, achieving strong agreement with human judgments; however, their computational cost and impracticality for processing sensitive data at scale limit real-world deployment. We address this gap by distilling the privacy assessment capabilities of Mistral Large 3 (675B) into lightweight encoder models with as few as 150M parameters. Leveraging a large-scale dataset of privacy-annotated texts spanning 10 diverse domains, we train efficient classifiers that preserve strong agreement with human annotations while dramatically reducing computational requirements. We validate our approach on human-annotated test data and demonstrate its practical utility as an evaluation metric for de-identification systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_29497 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Distilling Human-Aligned Privacy Sensitivity Assessment from Large Language Models Loiseau, Gabriel Sileo, Damien Riquet, Damien Meyer, Maxime Tommasi, Marc Computation and Language Accurate privacy evaluation of textual data remains a critical challenge in privacy-preserving natural language processing. Recent work has shown that large language models (LLMs) can serve as reliable privacy evaluators, achieving strong agreement with human judgments; however, their computational cost and impracticality for processing sensitive data at scale limit real-world deployment. We address this gap by distilling the privacy assessment capabilities of Mistral Large 3 (675B) into lightweight encoder models with as few as 150M parameters. Leveraging a large-scale dataset of privacy-annotated texts spanning 10 diverse domains, we train efficient classifiers that preserve strong agreement with human annotations while dramatically reducing computational requirements. We validate our approach on human-annotated test data and demonstrate its practical utility as an evaluation metric for de-identification systems. |
| title | Distilling Human-Aligned Privacy Sensitivity Assessment from Large Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2603.29497 |