DOREMI: Optimizing Long Tail Predictions in Document-Level Relation Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Menotti, Laura, Marchesin, Stefano, Silvello, Gianmaria
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917206366879744
author Menotti, Laura
Marchesin, Stefano
Silvello, Gianmaria
author_facet Menotti, Laura
Marchesin, Stefano
Silvello, Gianmaria
contents Document-Level Relation Extraction (DocRE) presents significant challenges due to its reliance on cross-sentence context and the long-tail distribution of relation types, where many relations have scarce training examples. In this work, we introduce DOcument-level Relation Extraction optiMizing the long taIl (DOREMI), an iterative framework that enhances underrepresented relations through minimal yet targeted manual annotations. Unlike previous approaches that rely on large-scale noisy data or heuristic denoising, DOREMI actively selects the most informative examples to improve training efficiency and robustness. DOREMI can be applied to any existing DocRE model and is effective at mitigating long-tail biases, offering a scalable solution to improve generalization on rare relations.
format Preprint
id arxiv_https___arxiv_org_abs_2601_11190
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DOREMI: Optimizing Long Tail Predictions in Document-Level Relation Extraction
Menotti, Laura
Marchesin, Stefano
Silvello, Gianmaria
Computation and Language
Document-Level Relation Extraction (DocRE) presents significant challenges due to its reliance on cross-sentence context and the long-tail distribution of relation types, where many relations have scarce training examples. In this work, we introduce DOcument-level Relation Extraction optiMizing the long taIl (DOREMI), an iterative framework that enhances underrepresented relations through minimal yet targeted manual annotations. Unlike previous approaches that rely on large-scale noisy data or heuristic denoising, DOREMI actively selects the most informative examples to improve training efficiency and robustness. DOREMI can be applied to any existing DocRE model and is effective at mitigating long-tail biases, offering a scalable solution to improve generalization on rare relations.
title DOREMI: Optimizing Long Tail Predictions in Document-Level Relation Extraction
topic Computation and Language
url https://arxiv.org/abs/2601.11190