GiusBERTo: A Legal Language Model for Personal Data De-identification in Italian Court of Auditors Decisions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Salierno, Giulio, Bertè, Rosamaria, Attias, Luca, Morrone, Carla, Pettazzoni, Dario, Battisti, Daniela
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929394089459712
author Salierno, Giulio
Bertè, Rosamaria
Attias, Luca
Morrone, Carla
Pettazzoni, Dario
Battisti, Daniela
author_facet Salierno, Giulio
Bertè, Rosamaria
Attias, Luca
Morrone, Carla
Pettazzoni, Dario
Battisti, Daniela
contents Recent advances in Natural Language Processing have demonstrated the effectiveness of pretrained language models like BERT for a variety of downstream tasks. We present GiusBERTo, the first BERT-based model specialized for anonymizing personal data in Italian legal documents. GiusBERTo is trained on a large dataset of Court of Auditors decisions to recognize entities to anonymize, including names, dates, locations, while retaining contextual relevance. We evaluate GiusBERTo on a held-out test set and achieve 97% token-level accuracy. GiusBERTo provides the Italian legal community with an accurate and tailored BERT model for de-identification, balancing privacy and data protection.
format Preprint
id arxiv_https___arxiv_org_abs_2406_15032
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GiusBERTo: A Legal Language Model for Personal Data De-identification in Italian Court of Auditors Decisions
Salierno, Giulio
Bertè, Rosamaria
Attias, Luca
Morrone, Carla
Pettazzoni, Dario
Battisti, Daniela
Computation and Language
Artificial Intelligence
Recent advances in Natural Language Processing have demonstrated the effectiveness of pretrained language models like BERT for a variety of downstream tasks. We present GiusBERTo, the first BERT-based model specialized for anonymizing personal data in Italian legal documents. GiusBERTo is trained on a large dataset of Court of Auditors decisions to recognize entities to anonymize, including names, dates, locations, while retaining contextual relevance. We evaluate GiusBERTo on a held-out test set and achieve 97% token-level accuracy. GiusBERTo provides the Italian legal community with an accurate and tailored BERT model for de-identification, balancing privacy and data protection.
title GiusBERTo: A Legal Language Model for Personal Data De-identification in Italian Court of Auditors Decisions
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2406.15032