GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zaratiana, Urchade, Lewis, Ash, Hurn-Maloney, George
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917479761051648
author Zaratiana, Urchade
Lewis, Ash
Hurn-Maloney, George
author_facet Zaratiana, Urchade
Lewis, Ash
Hurn-Maloney, George
contents Reliable detection of personally identifiable information (PII) is increasingly important across modern data-processing systems, yet the task remains difficult: PII spans are heterogeneous, locale-dependent, context-sensitive, and often embedded in noisy or semi-structured documents. We present GLiNER2-PII, a small 0.3B-parameter model adapted from GLiNER2 and designed to recognize a broad taxonomy of 42 PII entity types at character-span resolution. Training such systems, however, is constrained by the scarcity of shareable annotated data and the privacy risks associated with collecting real PII at scale. To address this challenge, we construct a multilingual synthetic corpus of 4,910 annotated texts using a constraint-driven generation pipeline that produces diverse, realistic examples across languages, domains, formats, and entity distributions. On the challenging SPY benchmark, GLiNER2-PII achieves the highest span-level F1 among five compared systems, including OpenAI Privacy Filter and three GLiNER-based detectors. We publicly release the model on Hugging Face to support further research and practical deployment of open PII detection systems.
format Preprint
id arxiv_https___arxiv_org_abs_2605_09973
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction
Zaratiana, Urchade
Lewis, Ash
Hurn-Maloney, George
Computation and Language
Artificial Intelligence
Reliable detection of personally identifiable information (PII) is increasingly important across modern data-processing systems, yet the task remains difficult: PII spans are heterogeneous, locale-dependent, context-sensitive, and often embedded in noisy or semi-structured documents. We present GLiNER2-PII, a small 0.3B-parameter model adapted from GLiNER2 and designed to recognize a broad taxonomy of 42 PII entity types at character-span resolution. Training such systems, however, is constrained by the scarcity of shareable annotated data and the privacy risks associated with collecting real PII at scale. To address this challenge, we construct a multilingual synthetic corpus of 4,910 annotated texts using a constraint-driven generation pipeline that produces diverse, realistic examples across languages, domains, formats, and entity distributions. On the challenging SPY benchmark, GLiNER2-PII achieves the highest span-level F1 among five compared systems, including OpenAI Privacy Filter and three GLiNER-based detectors. We publicly release the model on Hugging Face to support further research and practical deployment of open PII detection systems.
title GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.09973