Do not be greedy, Think Twice: Sampling and Selection for Document-level Information Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zubillaga, Mikel, Sainz, Oscar, de Lacalle, Oier Lopez, Agirre, Eneko
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916060117073920
author Zubillaga, Mikel
Sainz, Oscar
de Lacalle, Oier Lopez
Agirre, Eneko
author_facet Zubillaga, Mikel
Sainz, Oscar
de Lacalle, Oier Lopez
Agirre, Eneko
contents Document-level Information Extraction (DocIE) aims to produce an output template with the entities, relations, and events of interest occurring in the given document. Standard practices include prompting decoder-only LLMs using greedy decoding to avoid output variability. Rather than treating this variability as a limitation, we show that sampling can produce substantially better solutions than greedy decoding, especially when using reasoning models. We thus propose ThinkTwice, a sampling and selection framework in which the LLM generates multiple candidate templates for a given document, and a selection module chooses the most suitable one. We introduce both an unsupervised method that exploits agreement across generated outputs, and a supervised selection method using reward models trained on labeled DocIE data. To address the scarcity of golden reasoning trajectories for DocIE, we propose a rejection-sampling-based method to generate silver training data that pairs output templates with reasoning traces. Our experiments show the validity of unsupervised and supervised ThinkTwice, consistently outperforming greedy baselines and the supervised state-of-the-art.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18395
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Do not be greedy, Think Twice: Sampling and Selection for Document-level Information Extraction
Zubillaga, Mikel
Sainz, Oscar
de Lacalle, Oier Lopez
Agirre, Eneko
Computation and Language
Document-level Information Extraction (DocIE) aims to produce an output template with the entities, relations, and events of interest occurring in the given document. Standard practices include prompting decoder-only LLMs using greedy decoding to avoid output variability. Rather than treating this variability as a limitation, we show that sampling can produce substantially better solutions than greedy decoding, especially when using reasoning models. We thus propose ThinkTwice, a sampling and selection framework in which the LLM generates multiple candidate templates for a given document, and a selection module chooses the most suitable one. We introduce both an unsupervised method that exploits agreement across generated outputs, and a supervised selection method using reward models trained on labeled DocIE data. To address the scarcity of golden reasoning trajectories for DocIE, we propose a rejection-sampling-based method to generate silver training data that pairs output templates with reasoning traces. Our experiments show the validity of unsupervised and supervised ThinkTwice, consistently outperforming greedy baselines and the supervised state-of-the-art.
title Do not be greedy, Think Twice: Sampling and Selection for Document-level Information Extraction
topic Computation and Language
url https://arxiv.org/abs/2601.18395