HiligayNER: A Baseline Named Entity Recognition Model for Hiligaynon

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Teves, James Ald, Cal, Ray Daniel, Villaluz, Josh Magdiel, Malolos, Jean, Magtira, Mico, Rodriguez, Ramon, Abisado, Mideth, Imperial, Joseph Marvin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915548954099712
author Teves, James Ald
Cal, Ray Daniel
Villaluz, Josh Magdiel
Malolos, Jean
Magtira, Mico
Rodriguez, Ramon
Abisado, Mideth
Imperial, Joseph Marvin
author_facet Teves, James Ald
Cal, Ray Daniel
Villaluz, Josh Magdiel
Malolos, Jean
Magtira, Mico
Rodriguez, Ramon
Abisado, Mideth
Imperial, Joseph Marvin
contents The language of Hiligaynon, spoken predominantly by the people of Panay Island, Negros Occidental, and Soccsksargen in the Philippines, remains underrepresented in language processing research due to the absence of annotated corpora and baseline models. This study introduces HiligayNER, the first publicly available baseline model for the task of Named Entity Recognition (NER) in Hiligaynon. The dataset used to build HiligayNER contains over 8,000 annotated sentences collected from publicly available news articles, social media posts, and literary texts. Two Transformer-based models, mBERT and XLM-RoBERTa, were fine-tuned on this collected corpus to build versions of HiligayNER. Evaluation results show strong performance, with both models achieving over 80% in precision, recall, and F1-score across entity types. Furthermore, cross-lingual evaluation with Cebuano and Tagalog demonstrates promising transferability, suggesting the broader applicability of HiligayNER for multilingual NLP in low-resource settings. This work aims to contribute to language technology development for underrepresented Philippine languages, specifically for Hiligaynon, and support future research in regional language processing.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10776
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HiligayNER: A Baseline Named Entity Recognition Model for Hiligaynon
Teves, James Ald
Cal, Ray Daniel
Villaluz, Josh Magdiel
Malolos, Jean
Magtira, Mico
Rodriguez, Ramon
Abisado, Mideth
Imperial, Joseph Marvin
Computation and Language
The language of Hiligaynon, spoken predominantly by the people of Panay Island, Negros Occidental, and Soccsksargen in the Philippines, remains underrepresented in language processing research due to the absence of annotated corpora and baseline models. This study introduces HiligayNER, the first publicly available baseline model for the task of Named Entity Recognition (NER) in Hiligaynon. The dataset used to build HiligayNER contains over 8,000 annotated sentences collected from publicly available news articles, social media posts, and literary texts. Two Transformer-based models, mBERT and XLM-RoBERTa, were fine-tuned on this collected corpus to build versions of HiligayNER. Evaluation results show strong performance, with both models achieving over 80% in precision, recall, and F1-score across entity types. Furthermore, cross-lingual evaluation with Cebuano and Tagalog demonstrates promising transferability, suggesting the broader applicability of HiligayNER for multilingual NLP in low-resource settings. This work aims to contribute to language technology development for underrepresented Philippine languages, specifically for Hiligaynon, and support future research in regional language processing.
title HiligayNER: A Baseline Named Entity Recognition Model for Hiligaynon
topic Computation and Language
url https://arxiv.org/abs/2510.10776