Evaluating Large Language Models on Multiword Expressions in Multilingual and Code-Switched Contexts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: De Leon, Frances Laureano, Madabushi, Harish Tayyar, Lee, Mark G.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913812001587200
author De Leon, Frances Laureano
Madabushi, Harish Tayyar
Lee, Mark G.
author_facet De Leon, Frances Laureano
Madabushi, Harish Tayyar
Lee, Mark G.
contents Multiword expressions, characterised by non-compositional meanings and syntactic irregularities, are an example of nuanced language. These expressions can be used literally or idiomatically, leading to significant changes in meaning. While large language models have demonstrated strong performance across many tasks, their ability to handle such linguistic subtleties remains uncertain. Therefore, this study evaluates how state-of-the-art language models process the ambiguity of potentially idiomatic multiword expressions, particularly in contexts that are less frequent, where models are less likely to rely on memorisation. By evaluating models across in Portuguese and Galician, in addition to English, and using a novel code-switched dataset and a novel task, we find that large language models, despite their strengths, struggle with nuanced language. In particular, we find that the latest models, including GPT-4, fail to outperform the xlm-roBERTa-base baselines in both detection and semantic tasks, with especially poor performance on the novel tasks we introduce, despite its similarity to existing tasks. Overall, our results demonstrate that multiword expressions, especially those which are ambiguous, continue to be a challenge to models.
format Preprint
id arxiv_https___arxiv_org_abs_2504_20051
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Large Language Models on Multiword Expressions in Multilingual and Code-Switched Contexts
De Leon, Frances Laureano
Madabushi, Harish Tayyar
Lee, Mark G.
Computation and Language
Multiword expressions, characterised by non-compositional meanings and syntactic irregularities, are an example of nuanced language. These expressions can be used literally or idiomatically, leading to significant changes in meaning. While large language models have demonstrated strong performance across many tasks, their ability to handle such linguistic subtleties remains uncertain. Therefore, this study evaluates how state-of-the-art language models process the ambiguity of potentially idiomatic multiword expressions, particularly in contexts that are less frequent, where models are less likely to rely on memorisation. By evaluating models across in Portuguese and Galician, in addition to English, and using a novel code-switched dataset and a novel task, we find that large language models, despite their strengths, struggle with nuanced language. In particular, we find that the latest models, including GPT-4, fail to outperform the xlm-roBERTa-base baselines in both detection and semantic tasks, with especially poor performance on the novel tasks we introduce, despite its similarity to existing tasks. Overall, our results demonstrate that multiword expressions, especially those which are ambiguous, continue to be a challenge to models.
title Evaluating Large Language Models on Multiword Expressions in Multilingual and Code-Switched Contexts
topic Computation and Language
url https://arxiv.org/abs/2504.20051