NERCat: Fine-Tuning for Enhanced Named Entity Recognition in Catalan

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ferreres, Guillem Cadevall, Sanz, Marc Serrano, Gámez, Marc Bardeli, Basullas, Pol Gerdt, Ruiz, Francesc Tarres, Ferrero, Raul Quijada
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917959751958528
author Ferreres, Guillem Cadevall
Sanz, Marc Serrano
Gámez, Marc Bardeli
Basullas, Pol Gerdt
Ruiz, Francesc Tarres
Ferrero, Raul Quijada
author_facet Ferreres, Guillem Cadevall
Sanz, Marc Serrano
Gámez, Marc Bardeli
Basullas, Pol Gerdt
Ruiz, Francesc Tarres
Ferrero, Raul Quijada
contents Named Entity Recognition (NER) is a critical component of Natural Language Processing (NLP) for extracting structured information from unstructured text. However, for low-resource languages like Catalan, the performance of NER systems often suffers due to the lack of high-quality annotated datasets. This paper introduces NERCat, a fine-tuned version of the GLiNER[1] model, designed to improve NER performance specifically for Catalan text. We used a dataset of manually annotated Catalan television transcriptions to train and fine-tune the model, focusing on domains such as politics, sports, and culture. The evaluation results show significant improvements in precision, recall, and F1-score, particularly for underrepresented named entity categories such as Law, Product, and Facility. This study demonstrates the effectiveness of domain-specific fine-tuning in low-resource languages and highlights the potential for enhancing Catalan NLP applications through manual annotation and high-quality datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14173
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NERCat: Fine-Tuning for Enhanced Named Entity Recognition in Catalan
Ferreres, Guillem Cadevall
Sanz, Marc Serrano
Gámez, Marc Bardeli
Basullas, Pol Gerdt
Ruiz, Francesc Tarres
Ferrero, Raul Quijada
Computation and Language
68T50
I.2.7
Named Entity Recognition (NER) is a critical component of Natural Language Processing (NLP) for extracting structured information from unstructured text. However, for low-resource languages like Catalan, the performance of NER systems often suffers due to the lack of high-quality annotated datasets. This paper introduces NERCat, a fine-tuned version of the GLiNER[1] model, designed to improve NER performance specifically for Catalan text. We used a dataset of manually annotated Catalan television transcriptions to train and fine-tune the model, focusing on domains such as politics, sports, and culture. The evaluation results show significant improvements in precision, recall, and F1-score, particularly for underrepresented named entity categories such as Law, Product, and Facility. This study demonstrates the effectiveness of domain-specific fine-tuning in low-resource languages and highlights the potential for enhancing Catalan NLP applications through manual annotation and high-quality datasets.
title NERCat: Fine-Tuning for Enhanced Named Entity Recognition in Catalan
topic Computation and Language
68T50
I.2.7
url https://arxiv.org/abs/2503.14173