Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Golde, Jonas, Haller, Patrick, Hamborg, Felix, Risch, Julian, Akbik, Alan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916112948527104
author Golde, Jonas
Haller, Patrick
Hamborg, Felix
Risch, Julian
Akbik, Alan
author_facet Golde, Jonas
Haller, Patrick
Hamborg, Felix
Risch, Julian
Akbik, Alan
contents Most NLP tasks are modeled as supervised learning and thus require labeled training data to train effective models. However, manually producing such data at sufficient quality and quantity is known to be costly and time-intensive. Current research addresses this bottleneck by exploring a novel paradigm called zero-shot learning via dataset generation. Here, a powerful LLM is prompted with a task description to generate labeled data that can be used to train a downstream NLP model. For instance, an LLM might be prompted to "generate 500 movie reviews with positive overall sentiment, and another 500 with negative sentiment." The generated data could then be used to train a binary sentiment classifier, effectively leveraging an LLM as a teacher to a smaller student model. With this demo, we introduce Fabricator, an open-source Python toolkit for dataset generation. Fabricator implements common dataset generation workflows, supports a wide range of downstream NLP tasks (such as text classification, question answering, and entity recognition), and is integrated with well-known libraries to facilitate quick experimentation. With Fabricator, we aim to support researchers in conducting reproducible dataset generation experiments using LLMs and help practitioners apply this approach to train models for downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2309_09582
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs
Golde, Jonas
Haller, Patrick
Hamborg, Felix
Risch, Julian
Akbik, Alan
Computation and Language
Artificial Intelligence
Most NLP tasks are modeled as supervised learning and thus require labeled training data to train effective models. However, manually producing such data at sufficient quality and quantity is known to be costly and time-intensive. Current research addresses this bottleneck by exploring a novel paradigm called zero-shot learning via dataset generation. Here, a powerful LLM is prompted with a task description to generate labeled data that can be used to train a downstream NLP model. For instance, an LLM might be prompted to "generate 500 movie reviews with positive overall sentiment, and another 500 with negative sentiment." The generated data could then be used to train a binary sentiment classifier, effectively leveraging an LLM as a teacher to a smaller student model. With this demo, we introduce Fabricator, an open-source Python toolkit for dataset generation. Fabricator implements common dataset generation workflows, supports a wide range of downstream NLP tasks (such as text classification, question answering, and entity recognition), and is integrated with well-known libraries to facilitate quick experimentation. With Fabricator, we aim to support researchers in conducting reproducible dataset generation experiments using LLMs and help practitioners apply this approach to train models for downstream tasks.
title Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2309.09582