Word Meanings in Transformer Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Grindrod, Jumbly, Grindrod, Peter
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912541475602432
author Grindrod, Jumbly
Grindrod, Peter
author_facet Grindrod, Jumbly
Grindrod, Peter
contents We investigate how word meanings are represented in the transformer language models. Specifically, we focus on whether transformer models employ something analogous to a lexical store - where each word has an entry that contains semantic information. To do this, we extracted the token embedding space of RoBERTa-base and k-means clustered it into 200 clusters. In our first study, we then manually inspected the resultant clusters to consider whether they are sensitive to semantic information. In our second study, we tested whether the clusters are sensitive to five psycholinguistic measures: valence, concreteness, iconicity, taboo, and age of acquisition. Overall, our findings were very positive - there is a wide variety of semantic information encoded within the token embedding space. This serves to rule out certain "meaning eliminativist" hypotheses about how transformer LLMs process semantic information.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12863
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Word Meanings in Transformer Language Models
Grindrod, Jumbly
Grindrod, Peter
Computation and Language
Artificial Intelligence
We investigate how word meanings are represented in the transformer language models. Specifically, we focus on whether transformer models employ something analogous to a lexical store - where each word has an entry that contains semantic information. To do this, we extracted the token embedding space of RoBERTa-base and k-means clustered it into 200 clusters. In our first study, we then manually inspected the resultant clusters to consider whether they are sensitive to semantic information. In our second study, we tested whether the clusters are sensitive to five psycholinguistic measures: valence, concreteness, iconicity, taboo, and age of acquisition. Overall, our findings were very positive - there is a wide variety of semantic information encoded within the token embedding space. This serves to rule out certain "meaning eliminativist" hypotheses about how transformer LLMs process semantic information.
title Word Meanings in Transformer Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.12863