Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Basmov, Victoria, Goldberg, Yoav, Tsarfaty, Reut
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917636647944192
author Basmov, Victoria
Goldberg, Yoav
Tsarfaty, Reut
author_facet Basmov, Victoria
Goldberg, Yoav
Tsarfaty, Reut
contents We evaluate LLMs' language understanding capacities on simple inference tasks that most humans find trivial. Specifically, we target (i) grammatically-specified entailments, (ii) premises with evidential adverbs of uncertainty, and (iii) monotonicity entailments. We design evaluation sets for these tasks and conduct experiments in both zero-shot and chain-of-thought setups, and with multiple prompts and LLMs. The models exhibit moderate to low performance on these evaluation sets. Subsequent experiments show that embedding the premise in syntactic constructions that should preserve the entailment relations (presupposition triggers) or change them (non-factives), further confuses the models, causing them to either under-predict or over-predict certain entailment labels regardless of the true relation, and often disregarding the nature of the embedding context. Overall these results suggest that, despite LLMs' celebrated language understanding capacity, even the strongest models have blindspots with respect to certain types of entailments, and certain information-packaging structures act as ``blinds'' overshadowing the semantics of the embedded premise.
format Preprint
id arxiv_https___arxiv_org_abs_2305_14785
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds
Basmov, Victoria
Goldberg, Yoav
Tsarfaty, Reut
Computation and Language
Artificial Intelligence
We evaluate LLMs' language understanding capacities on simple inference tasks that most humans find trivial. Specifically, we target (i) grammatically-specified entailments, (ii) premises with evidential adverbs of uncertainty, and (iii) monotonicity entailments. We design evaluation sets for these tasks and conduct experiments in both zero-shot and chain-of-thought setups, and with multiple prompts and LLMs. The models exhibit moderate to low performance on these evaluation sets. Subsequent experiments show that embedding the premise in syntactic constructions that should preserve the entailment relations (presupposition triggers) or change them (non-factives), further confuses the models, causing them to either under-predict or over-predict certain entailment labels regardless of the true relation, and often disregarding the nature of the embedding context. Overall these results suggest that, despite LLMs' celebrated language understanding capacity, even the strongest models have blindspots with respect to certain types of entailments, and certain information-packaging structures act as ``blinds'' overshadowing the semantics of the embedded premise.
title Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2305.14785