Saved in:
Bibliographic Details
Main Authors: Miletić, Filip, Walde, Sabine Schulte im
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2401.15393
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929226145333248
author Miletić, Filip
Walde, Sabine Schulte im
author_facet Miletić, Filip
Walde, Sabine Schulte im
contents Multiword expressions (MWEs) are composed of multiple words and exhibit variable degrees of compositionality. As such, their meanings are notoriously difficult to model, and it is unclear to what extent this issue affects transformer architectures. Addressing this gap, we provide the first in-depth survey of MWE processing with transformer models. We overall find that they capture MWE semantics inconsistently, as shown by reliance on surface patterns and memorized information. MWE meaning is also strongly localized, predominantly in early layers of the architecture. Representations benefit from specific linguistic properties, such as lower semantic idiosyncrasy and ambiguity of target expressions. Our findings overall question the ability of transformer models to robustly capture fine-grained semantics. Furthermore, we highlight the need for more directly comparable evaluation setups.
format Preprint
id arxiv_https___arxiv_org_abs_2401_15393
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Semantics of Multiword Expressions in Transformer-Based Models: A Survey
Miletić, Filip
Walde, Sabine Schulte im
Computation and Language
Multiword expressions (MWEs) are composed of multiple words and exhibit variable degrees of compositionality. As such, their meanings are notoriously difficult to model, and it is unclear to what extent this issue affects transformer architectures. Addressing this gap, we provide the first in-depth survey of MWE processing with transformer models. We overall find that they capture MWE semantics inconsistently, as shown by reliance on surface patterns and memorized information. MWE meaning is also strongly localized, predominantly in early layers of the architecture. Representations benefit from specific linguistic properties, such as lower semantic idiosyncrasy and ambiguity of target expressions. Our findings overall question the ability of transformer models to robustly capture fine-grained semantics. Furthermore, we highlight the need for more directly comparable evaluation setups.
title Semantics of Multiword Expressions in Transformer-Based Models: A Survey
topic Computation and Language
url https://arxiv.org/abs/2401.15393