Exploring Interpretability of Independent Components of Word Embeddings with Automated Word Intruder Test

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Musil, Tomáš, Mareček, David
Natura: Preprint
Pubblicazione: 2022
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909303702552576
author Musil, Tomáš
Mareček, David
author_facet Musil, Tomáš
Mareček, David
contents Independent Component Analysis (ICA) is an algorithm originally developed for finding separate sources in a mixed signal, such as a recording of multiple people in the same room speaking at the same time. Unlike Principal Component Analysis (PCA), ICA permits the representation of a word as an unstructured set of features, without any particular feature being deemed more significant than the others. In this paper, we used ICA to analyze word embeddings. We have found that ICA can be used to find semantic features of the words, and these features can easily be combined to search for words that satisfy the combination. We show that most of the independent components represent such features. To quantify the interpretability of the components, we use the word intruder test, performed both by humans and by large language models. We propose to use the automated version of the word intruder test as a fast and inexpensive way of quantifying vector interpretability without the need for human effort.
format Preprint
id arxiv_https___arxiv_org_abs_2212_09580
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Exploring Interpretability of Independent Components of Word Embeddings with Automated Word Intruder Test
Musil, Tomáš
Mareček, David
Computation and Language
Independent Component Analysis (ICA) is an algorithm originally developed for finding separate sources in a mixed signal, such as a recording of multiple people in the same room speaking at the same time. Unlike Principal Component Analysis (PCA), ICA permits the representation of a word as an unstructured set of features, without any particular feature being deemed more significant than the others. In this paper, we used ICA to analyze word embeddings. We have found that ICA can be used to find semantic features of the words, and these features can easily be combined to search for words that satisfy the combination. We show that most of the independent components represent such features. To quantify the interpretability of the components, we use the word intruder test, performed both by humans and by large language models. We propose to use the automated version of the word intruder test as a fast and inexpensive way of quantifying vector interpretability without the need for human effort.
title Exploring Interpretability of Independent Components of Word Embeddings with Automated Word Intruder Test
topic Computation and Language
url https://arxiv.org/abs/2212.09580