Rethinking STS and NLI in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yuxia, Wang, Minghan, Nakov, Preslav
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916113894342656
author Wang, Yuxia
Wang, Minghan
Nakov, Preslav
author_facet Wang, Yuxia
Wang, Minghan
Nakov, Preslav
contents Recent years have seen the rise of large language models (LLMs), where practitioners use task-specific prompts; this was shown to be effective for a variety of tasks. However, when applied to semantic textual similarity (STS) and natural language inference (NLI), the effectiveness of LLMs turns out to be limited by low-resource domain accuracy, model overconfidence, and difficulty to capture the disagreements between human judgements. With this in mind, here we try to rethink STS and NLI in the era of LLMs. We first evaluate the performance of STS and NLI in the clinical/biomedical domain, and then we assess LLMs' predictive confidence and their capability of capturing collective human opinions. We find that these old problems are still to be properly addressed in the era of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2309_08969
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Rethinking STS and NLI in Large Language Models
Wang, Yuxia
Wang, Minghan
Nakov, Preslav
Computation and Language
Recent years have seen the rise of large language models (LLMs), where practitioners use task-specific prompts; this was shown to be effective for a variety of tasks. However, when applied to semantic textual similarity (STS) and natural language inference (NLI), the effectiveness of LLMs turns out to be limited by low-resource domain accuracy, model overconfidence, and difficulty to capture the disagreements between human judgements. With this in mind, here we try to rethink STS and NLI in the era of LLMs. We first evaluate the performance of STS and NLI in the clinical/biomedical domain, and then we assess LLMs' predictive confidence and their capability of capturing collective human opinions. We find that these old problems are still to be properly addressed in the era of LLMs.
title Rethinking STS and NLI in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2309.08969