ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Feuer, Benjamin, Liu, Yurong, Hegde, Chinmay, Freire, Juliana
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916361595256832
author Feuer, Benjamin
Liu, Yurong
Hegde, Chinmay
Freire, Juliana
author_facet Feuer, Benjamin
Liu, Yurong
Hegde, Chinmay
Freire, Juliana
contents Existing deep-learning approaches to semantic column type annotation (CTA) have important shortcomings: they rely on semantic types which are fixed at training time; require a large number of training samples per type and incur large run-time inference costs; and their performance can degrade when evaluated on novel datasets, even when types remain constant. Large language models have exhibited strong zero-shot classification performance on a wide range of tasks and in this paper we explore their use for CTA. We introduce ArcheType, a simple, practical method for context sampling, prompt serialization, model querying, and label remapping, which enables large language models to solve CTA problems in a fully zero-shot manner. We ablate each component of our method separately, and establish that improvements to context sampling and label remapping provide the most consistent gains. ArcheType establishes a new state-of-the-art performance on zero-shot CTA benchmarks (including three new domain-specific benchmarks which we release along with this paper), and when used in conjunction with classical CTA techniques, it outperforms a SOTA DoDuo model on the fine-tuned SOTAB benchmark. Our code is available at https://github.com/penfever/ArcheType.
format Preprint
id arxiv_https___arxiv_org_abs_2310_18208
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models
Feuer, Benjamin
Liu, Yurong
Hegde, Chinmay
Freire, Juliana
Computation and Language
Machine Learning
H.3.3; H.3; I.2; I.2.7
Existing deep-learning approaches to semantic column type annotation (CTA) have important shortcomings: they rely on semantic types which are fixed at training time; require a large number of training samples per type and incur large run-time inference costs; and their performance can degrade when evaluated on novel datasets, even when types remain constant. Large language models have exhibited strong zero-shot classification performance on a wide range of tasks and in this paper we explore their use for CTA. We introduce ArcheType, a simple, practical method for context sampling, prompt serialization, model querying, and label remapping, which enables large language models to solve CTA problems in a fully zero-shot manner. We ablate each component of our method separately, and establish that improvements to context sampling and label remapping provide the most consistent gains. ArcheType establishes a new state-of-the-art performance on zero-shot CTA benchmarks (including three new domain-specific benchmarks which we release along with this paper), and when used in conjunction with classical CTA techniques, it outperforms a SOTA DoDuo model on the fine-tuned SOTAB benchmark. Our code is available at https://github.com/penfever/ArcheType.
title ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models
topic Computation and Language
Machine Learning
H.3.3; H.3; I.2; I.2.7
url https://arxiv.org/abs/2310.18208