Text clustering applied to data augmentation in legal contexts

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Freitas, Lucas José Gonçalves, Rodrigues, Thaís, Rodrigues, Guilherme, Edokawa, Pamella, Farias, Ariane
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916204325634048
author Freitas, Lucas José Gonçalves
Rodrigues, Thaís
Rodrigues, Guilherme
Edokawa, Pamella
Farias, Ariane
author_facet Freitas, Lucas José Gonçalves
Rodrigues, Thaís
Rodrigues, Guilherme
Edokawa, Pamella
Farias, Ariane
contents Data analysis and machine learning are of preeminent importance in the legal domain, especially in tasks like clustering and text classification. In this study, we harnessed the power of natural language processing tools to enhance datasets meticulously curated by experts. This process significantly improved the classification workflow for legal texts using machine learning techniques. We considered the Sustainable Development Goals (SDGs) data from the United Nations 2030 Agenda as a practical case study. Data augmentation clustering-based strategy led to remarkable enhancements in the accuracy and sensitivity metrics of classification models. For certain SDGs within the 2030 Agenda, we observed performance gains of over 15%. In some cases, the example base expanded by a noteworthy factor of 5. When dealing with unclassified legal texts, data augmentation strategies centered around clustering prove to be highly effective. They provide a valuable means to expand the existing knowledge base without the need for labor-intensive manual classification efforts.
format Preprint
id arxiv_https___arxiv_org_abs_2404_08683
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Text clustering applied to data augmentation in legal contexts
Freitas, Lucas José Gonçalves
Rodrigues, Thaís
Rodrigues, Guilherme
Edokawa, Pamella
Farias, Ariane
Computation and Language
Machine Learning
Data analysis and machine learning are of preeminent importance in the legal domain, especially in tasks like clustering and text classification. In this study, we harnessed the power of natural language processing tools to enhance datasets meticulously curated by experts. This process significantly improved the classification workflow for legal texts using machine learning techniques. We considered the Sustainable Development Goals (SDGs) data from the United Nations 2030 Agenda as a practical case study. Data augmentation clustering-based strategy led to remarkable enhancements in the accuracy and sensitivity metrics of classification models. For certain SDGs within the 2030 Agenda, we observed performance gains of over 15%. In some cases, the example base expanded by a noteworthy factor of 5. When dealing with unclassified legal texts, data augmentation strategies centered around clustering prove to be highly effective. They provide a valuable means to expand the existing knowledge base without the need for labor-intensive manual classification efforts.
title Text clustering applied to data augmentation in legal contexts
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2404.08683