Is There a One-Model-Fits-All Approach to Information Extraction? Revisiting Task Definition Biases

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Huang, Wenhao, He, Qianyu, Li, Zhixu, Liang, Jiaqing, Xiao, Yanghua
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917621598781440
author Huang, Wenhao
He, Qianyu
Li, Zhixu
Liang, Jiaqing
Xiao, Yanghua
author_facet Huang, Wenhao
He, Qianyu
Li, Zhixu
Liang, Jiaqing
Xiao, Yanghua
contents Definition bias is a negative phenomenon that can mislead models. Definition bias in information extraction appears not only across datasets from different domains but also within datasets sharing the same domain. We identify two types of definition bias in IE: bias among information extraction datasets and bias between information extraction datasets and instruction tuning datasets. To systematically investigate definition bias, we conduct three probing experiments to quantitatively analyze it and discover the limitations of unified information extraction and large language models in solving definition bias. To mitigate definition bias in information extraction, we propose a multi-stage framework consisting of definition bias measurement, bias-aware fine-tuning, and task-specific bias mitigation. Experimental results demonstrate the effectiveness of our framework in addressing definition bias. Resources of this paper can be found at https://github.com/EZ-hwh/definition-bias
format Preprint
id arxiv_https___arxiv_org_abs_2403_16396
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Is There a One-Model-Fits-All Approach to Information Extraction? Revisiting Task Definition Biases
Huang, Wenhao
He, Qianyu
Li, Zhixu
Liang, Jiaqing
Xiao, Yanghua
Computation and Language
Definition bias is a negative phenomenon that can mislead models. Definition bias in information extraction appears not only across datasets from different domains but also within datasets sharing the same domain. We identify two types of definition bias in IE: bias among information extraction datasets and bias between information extraction datasets and instruction tuning datasets. To systematically investigate definition bias, we conduct three probing experiments to quantitatively analyze it and discover the limitations of unified information extraction and large language models in solving definition bias. To mitigate definition bias in information extraction, we propose a multi-stage framework consisting of definition bias measurement, bias-aware fine-tuning, and task-specific bias mitigation. Experimental results demonstrate the effectiveness of our framework in addressing definition bias. Resources of this paper can be found at https://github.com/EZ-hwh/definition-bias
title Is There a One-Model-Fits-All Approach to Information Extraction? Revisiting Task Definition Biases
topic Computation and Language
url https://arxiv.org/abs/2403.16396