Can Large Language Models Follow Concept Annotation Guidelines? A Case Study on Scientific and Financial Domains

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fonseca, Marcio, Cohen, Shay B.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914849563344896
author Fonseca, Marcio
Cohen, Shay B.
author_facet Fonseca, Marcio
Cohen, Shay B.
contents Although large language models (LLMs) exhibit remarkable capacity to leverage in-context demonstrations, it is still unclear to what extent they can learn new concepts or facts from ground-truth labels. To address this question, we examine the capacity of instruction-tuned LLMs to follow in-context concept guidelines for sentence labeling tasks. We design guidelines that present different types of factual and counterfactual concept definitions, which are used as prompts for zero-shot sentence classification tasks. Our results show that although concept definitions consistently help in task performance, only the larger models (with 70B parameters or more) have limited ability to work under counterfactual contexts. Importantly, only proprietary models such as GPT-3.5 and GPT-4 can recognize nonsensical guidelines, which we hypothesize is due to more sophisticated alignment methods. Finally, we find that Falcon-180B-chat is outperformed by Llama-2-70B-chat is most cases, which indicates that careful fine-tuning is more effective than increasing model scale. Altogether, our simple evaluation method reveals significant gaps in concept understanding between the most capable open-source language models and the leading proprietary APIs.
format Preprint
id arxiv_https___arxiv_org_abs_2311_08704
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Can Large Language Models Follow Concept Annotation Guidelines? A Case Study on Scientific and Financial Domains
Fonseca, Marcio
Cohen, Shay B.
Computation and Language
Artificial Intelligence
Although large language models (LLMs) exhibit remarkable capacity to leverage in-context demonstrations, it is still unclear to what extent they can learn new concepts or facts from ground-truth labels. To address this question, we examine the capacity of instruction-tuned LLMs to follow in-context concept guidelines for sentence labeling tasks. We design guidelines that present different types of factual and counterfactual concept definitions, which are used as prompts for zero-shot sentence classification tasks. Our results show that although concept definitions consistently help in task performance, only the larger models (with 70B parameters or more) have limited ability to work under counterfactual contexts. Importantly, only proprietary models such as GPT-3.5 and GPT-4 can recognize nonsensical guidelines, which we hypothesize is due to more sophisticated alignment methods. Finally, we find that Falcon-180B-chat is outperformed by Llama-2-70B-chat is most cases, which indicates that careful fine-tuning is more effective than increasing model scale. Altogether, our simple evaluation method reveals significant gaps in concept understanding between the most capable open-source language models and the leading proprietary APIs.
title Can Large Language Models Follow Concept Annotation Guidelines? A Case Study on Scientific and Financial Domains
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2311.08704