Challenges in Expanding Portuguese Resources: A View from Open Information Extraction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Souza, Marlo, Cabral, Bruno, Claro, Daniela, Salvador, Lais
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929683192348672
author Souza, Marlo
Cabral, Bruno
Claro, Daniela
Salvador, Lais
author_facet Souza, Marlo
Cabral, Bruno
Claro, Daniela
Salvador, Lais
contents Open Information Extraction (Open IE) is the task of extracting structured information from textual documents, independent of domain. While traditional Open IE methods were based on unsupervised approaches, recently, with the emergence of robust annotated datasets, new data-based approaches have been developed to achieve better results. These innovations, however, have focused mainly on the English language due to a lack of datasets and the difficulty of constructing such resources for other languages. In this work, we present a high-quality manually annotated corpus for Open Information Extraction in the Portuguese language, based on a rigorous methodology grounded in established semantic theories. We discuss the challenges encountered in the annotation process, propose a set of structural and contextual annotation rules, and validate our corpus by evaluating the performance of state-of-the-art Open IE systems. Our resource addresses the lack of datasets for Open IE in Portuguese and can support the development and evaluation of new methods and systems in this area.
format Preprint
id arxiv_https___arxiv_org_abs_2501_11851
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Challenges in Expanding Portuguese Resources: A View from Open Information Extraction
Souza, Marlo
Cabral, Bruno
Claro, Daniela
Salvador, Lais
Computation and Language
Open Information Extraction (Open IE) is the task of extracting structured information from textual documents, independent of domain. While traditional Open IE methods were based on unsupervised approaches, recently, with the emergence of robust annotated datasets, new data-based approaches have been developed to achieve better results. These innovations, however, have focused mainly on the English language due to a lack of datasets and the difficulty of constructing such resources for other languages. In this work, we present a high-quality manually annotated corpus for Open Information Extraction in the Portuguese language, based on a rigorous methodology grounded in established semantic theories. We discuss the challenges encountered in the annotation process, propose a set of structural and contextual annotation rules, and validate our corpus by evaluating the performance of state-of-the-art Open IE systems. Our resource addresses the lack of datasets for Open IE in Portuguese and can support the development and evaluation of new methods and systems in this area.
title Challenges in Expanding Portuguese Resources: A View from Open Information Extraction
topic Computation and Language
url https://arxiv.org/abs/2501.11851