PICLe: Pseudo-Annotations for In-Context Learning in Low-Resource Named Entity Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mamooler, Sepideh, Montariol, Syrielle, Mathis, Alexander, Bosselut, Antoine
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916669170909184
author Mamooler, Sepideh
Montariol, Syrielle
Mathis, Alexander
Bosselut, Antoine
author_facet Mamooler, Sepideh
Montariol, Syrielle
Mathis, Alexander
Bosselut, Antoine
contents In-context learning (ICL) enables Large Language Models (LLMs) to perform tasks using few demonstrations, facilitating task adaptation when labeled examples are hard to obtain. However, ICL is sensitive to the choice of demonstrations, and it remains unclear which demonstration attributes enable in-context generalization. In this work, we conduct a perturbation study of in-context demonstrations for low-resource Named Entity Detection (NED). Our surprising finding is that in-context demonstrations with partially correct annotated entity mentions can be as effective for task transfer as fully correct demonstrations. Based off our findings, we propose Pseudo-annotated In-Context Learning (PICLe), a framework for in-context learning with noisy, pseudo-annotated demonstrations. PICLe leverages LLMs to annotate many demonstrations in a zero-shot first pass. We then cluster these synthetic demonstrations, sample specific sets of in-context demonstrations from each cluster, and predict entity mentions using each set independently. Finally, we use self-verification to select the final set of entity mentions. We evaluate PICLe on five biomedical NED datasets and show that, with zero human annotation, PICLe outperforms ICL in low-resource settings where limited gold examples can be used as in-context demonstrations.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11923
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PICLe: Pseudo-Annotations for In-Context Learning in Low-Resource Named Entity Detection
Mamooler, Sepideh
Montariol, Syrielle
Mathis, Alexander
Bosselut, Antoine
Computation and Language
Artificial Intelligence
In-context learning (ICL) enables Large Language Models (LLMs) to perform tasks using few demonstrations, facilitating task adaptation when labeled examples are hard to obtain. However, ICL is sensitive to the choice of demonstrations, and it remains unclear which demonstration attributes enable in-context generalization. In this work, we conduct a perturbation study of in-context demonstrations for low-resource Named Entity Detection (NED). Our surprising finding is that in-context demonstrations with partially correct annotated entity mentions can be as effective for task transfer as fully correct demonstrations. Based off our findings, we propose Pseudo-annotated In-Context Learning (PICLe), a framework for in-context learning with noisy, pseudo-annotated demonstrations. PICLe leverages LLMs to annotate many demonstrations in a zero-shot first pass. We then cluster these synthetic demonstrations, sample specific sets of in-context demonstrations from each cluster, and predict entity mentions using each set independently. Finally, we use self-verification to select the final set of entity mentions. We evaluate PICLe on five biomedical NED datasets and show that, with zero human annotation, PICLe outperforms ICL in low-resource settings where limited gold examples can be used as in-context demonstrations.
title PICLe: Pseudo-Annotations for In-Context Learning in Low-Resource Named Entity Detection
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.11923