InstructIE: A Bilingual Instruction-based Information Extraction Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gui, Honghao, Qiao, Shuofei, Zhang, Jintian, Ye, Hongbin, Sun, Mengshu, Liang, Lei, Pan, Jeff Z., Chen, Huajun, Zhang, Ningyu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914890029989888
author Gui, Honghao
Qiao, Shuofei
Zhang, Jintian
Ye, Hongbin
Sun, Mengshu
Liang, Lei
Pan, Jeff Z.
Chen, Huajun
Zhang, Ningyu
author_facet Gui, Honghao
Qiao, Shuofei
Zhang, Jintian
Ye, Hongbin
Sun, Mengshu
Liang, Lei
Pan, Jeff Z.
Chen, Huajun
Zhang, Ningyu
contents Large language models can perform well on general natural language tasks, but their effectiveness is still suboptimal for information extraction (IE). Recent works indicate that the main reason lies in the lack of extensive data on IE instructions. Note that the existing datasets on IE instructions not only have limited coverage but also involve high construction costs. To address this issue, we introduce InstructIE, a bilingual instruction-based IE dataset, which covers 12 diverse domains. We propose KG2Instruction, a framework specifically for the automatic generation of such datasets. Additionally, we manually annotate the test set. Experimental results demonstrate that large language models trained with InstructIE can not only obtain better IE capabilities but also enhance zero-shot performance compared with baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2305_11527
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle InstructIE: A Bilingual Instruction-based Information Extraction Dataset
Gui, Honghao
Qiao, Shuofei
Zhang, Jintian
Ye, Hongbin
Sun, Mengshu
Liang, Lei
Pan, Jeff Z.
Chen, Huajun
Zhang, Ningyu
Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
Large language models can perform well on general natural language tasks, but their effectiveness is still suboptimal for information extraction (IE). Recent works indicate that the main reason lies in the lack of extensive data on IE instructions. Note that the existing datasets on IE instructions not only have limited coverage but also involve high construction costs. To address this issue, we introduce InstructIE, a bilingual instruction-based IE dataset, which covers 12 diverse domains. We propose KG2Instruction, a framework specifically for the automatic generation of such datasets. Additionally, we manually annotate the test set. Experimental results demonstrate that large language models trained with InstructIE can not only obtain better IE capabilities but also enhance zero-shot performance compared with baselines.
title InstructIE: A Bilingual Instruction-based Information Extraction Dataset
topic Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2305.11527