Revisiting subword tokenization: A case study on affixal negation in large language models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Truong, Thinh Hung, Otmakhova, Yulia, Verspoor, Karin, Cohn, Trevor, Baldwin, Timothy
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913298395430912
author Truong, Thinh Hung
Otmakhova, Yulia
Verspoor, Karin
Cohn, Trevor
Baldwin, Timothy
author_facet Truong, Thinh Hung
Otmakhova, Yulia
Verspoor, Karin
Cohn, Trevor
Baldwin, Timothy
contents In this work, we measure the impact of affixal negation on modern English large language models (LLMs). In affixal negation, the negated meaning is expressed through a negative morpheme, which is potentially challenging for LLMs as their tokenizers are often not morphologically plausible. We conduct extensive experiments using LLMs with different subword tokenization methods, which lead to several insights on the interaction between tokenization performance and negation sensitivity. Despite some interesting mismatches between tokenization accuracy and negation detection performance, we show that models can, on the whole, reliably recognize the meaning of affixal negation.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02421
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Revisiting subword tokenization: A case study on affixal negation in large language models
Truong, Thinh Hung
Otmakhova, Yulia
Verspoor, Karin
Cohn, Trevor
Baldwin, Timothy
Computation and Language
In this work, we measure the impact of affixal negation on modern English large language models (LLMs). In affixal negation, the negated meaning is expressed through a negative morpheme, which is potentially challenging for LLMs as their tokenizers are often not morphologically plausible. We conduct extensive experiments using LLMs with different subword tokenization methods, which lead to several insights on the interaction between tokenization performance and negation sensitivity. Despite some interesting mismatches between tokenization accuracy and negation detection performance, we show that models can, on the whole, reliably recognize the meaning of affixal negation.
title Revisiting subword tokenization: A case study on affixal negation in large language models
topic Computation and Language
url https://arxiv.org/abs/2404.02421