Adapting PromptORE for Modern History: Information Extraction from Hispanic Monarchy Documents of the XVIth Century

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hidalgo, Hèctor Loopez, Boeglin, Michel, Kahn, David, Mothe, Josiane, Ortiz, Diego, Panzoli, David
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913373138976768
author Hidalgo, Hèctor Loopez
Boeglin, Michel
Kahn, David
Mothe, Josiane
Ortiz, Diego
Panzoli, David
author_facet Hidalgo, Hèctor Loopez
Boeglin, Michel
Kahn, David
Mothe, Josiane
Ortiz, Diego
Panzoli, David
contents Semantic relations among entities are a widely accepted method for relation extraction. PromptORE (Prompt-based Open Relation Extraction) was designed to improve relation extraction with Large Language Models on generalistic documents. However, it is less effective when applied to historical documents, in languages other than English. In this study, we introduce an adaptation of PromptORE to extract relations from specialized documents, namely digital transcripts of trials from the Spanish Inquisition. Our approach involves fine-tuning transformer models with their pretraining objective on the data they will perform inference. We refer to this process as "biasing". Our Biased PromptORE addresses complex entity placements and genderism that occur in Spanish texts. We solve these issues by prompt engineering. We evaluate our method using Encoder-like models, corroborating our findings with experts' assessments. Additionally, we evaluate the performance using a binomial classification benchmark. Our results show a substantial improvement in accuracy -up to a 50% improvement with our Biased PromptORE models in comparison to the baseline models using standard PromptORE.
format Preprint
id arxiv_https___arxiv_org_abs_2406_00027
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adapting PromptORE for Modern History: Information Extraction from Hispanic Monarchy Documents of the XVIth Century
Hidalgo, Hèctor Loopez
Boeglin, Michel
Kahn, David
Mothe, Josiane
Ortiz, Diego
Panzoli, David
Computation and Language
Information Retrieval
Machine Learning
H.3; H.3.2; H.4; I.7; I.2; I.2.7
Semantic relations among entities are a widely accepted method for relation extraction. PromptORE (Prompt-based Open Relation Extraction) was designed to improve relation extraction with Large Language Models on generalistic documents. However, it is less effective when applied to historical documents, in languages other than English. In this study, we introduce an adaptation of PromptORE to extract relations from specialized documents, namely digital transcripts of trials from the Spanish Inquisition. Our approach involves fine-tuning transformer models with their pretraining objective on the data they will perform inference. We refer to this process as "biasing". Our Biased PromptORE addresses complex entity placements and genderism that occur in Spanish texts. We solve these issues by prompt engineering. We evaluate our method using Encoder-like models, corroborating our findings with experts' assessments. Additionally, we evaluate the performance using a binomial classification benchmark. Our results show a substantial improvement in accuracy -up to a 50% improvement with our Biased PromptORE models in comparison to the baseline models using standard PromptORE.
title Adapting PromptORE for Modern History: Information Extraction from Hispanic Monarchy Documents of the XVIth Century
topic Computation and Language
Information Retrieval
Machine Learning
H.3; H.3.2; H.4; I.7; I.2; I.2.7
url https://arxiv.org/abs/2406.00027