GuideX: Guided Synthetic Data Generation for Zero-Shot Information Extraction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: De La Fuente, Neil, Sainz, Oscar, García-Ferrero, Iker, Agirre, Eneko
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909630525865984
author De La Fuente, Neil
Sainz, Oscar
García-Ferrero, Iker
Agirre, Eneko
author_facet De La Fuente, Neil
Sainz, Oscar
García-Ferrero, Iker
Agirre, Eneko
contents Information Extraction (IE) systems are traditionally domain-specific, requiring costly adaptation that involves expert schema design, data annotation, and model training. While Large Language Models have shown promise in zero-shot IE, performance degrades significantly in unseen domains where label definitions differ. This paper introduces GUIDEX, a novel method that automatically defines domain-specific schemas, infers guidelines, and generates synthetically labeled instances, allowing for better out-of-domain generalization. Fine-tuning Llama 3.1 with GUIDEX sets a new state-of-the-art across seven zeroshot Named Entity Recognition benchmarks. Models trained with GUIDEX gain up to 7 F1 points over previous methods without humanlabeled data, and nearly 2 F1 points higher when combined with it. Models trained on GUIDEX demonstrate enhanced comprehension of complex, domain-specific annotation schemas. Code, models, and synthetic datasets are available at neilus03.github.io/guidex.com
format Preprint
id arxiv_https___arxiv_org_abs_2506_00649
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GuideX: Guided Synthetic Data Generation for Zero-Shot Information Extraction
De La Fuente, Neil
Sainz, Oscar
García-Ferrero, Iker
Agirre, Eneko
Computation and Language
Information Extraction (IE) systems are traditionally domain-specific, requiring costly adaptation that involves expert schema design, data annotation, and model training. While Large Language Models have shown promise in zero-shot IE, performance degrades significantly in unseen domains where label definitions differ. This paper introduces GUIDEX, a novel method that automatically defines domain-specific schemas, infers guidelines, and generates synthetically labeled instances, allowing for better out-of-domain generalization. Fine-tuning Llama 3.1 with GUIDEX sets a new state-of-the-art across seven zeroshot Named Entity Recognition benchmarks. Models trained with GUIDEX gain up to 7 F1 points over previous methods without humanlabeled data, and nearly 2 F1 points higher when combined with it. Models trained on GUIDEX demonstrate enhanced comprehension of complex, domain-specific annotation schemas. Code, models, and synthetic datasets are available at neilus03.github.io/guidex.com
title GuideX: Guided Synthetic Data Generation for Zero-Shot Information Extraction
topic Computation and Language
url https://arxiv.org/abs/2506.00649