Are Word Embedding Methods Stable and Should We Care About It?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Borah, Angana, Barman, Manash Pratim, Awekar, Amit
Format: Preprint
Veröffentlicht: 2021
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914831825633280
author Borah, Angana
Barman, Manash Pratim
Awekar, Amit
author_facet Borah, Angana
Barman, Manash Pratim
Awekar, Amit
contents A representation learning method is considered stable if it consistently generates similar representation of the given data across multiple runs. Word Embedding Methods (WEMs) are a class of representation learning methods that generate dense vector representation for each word in the given text data. The central idea of this paper is to explore the stability measurement of WEMs using intrinsic evaluation based on word similarity. We experiment with three popular WEMs: Word2Vec, GloVe, and fastText. For stability measurement, we investigate the effect of five parameters involved in training these models. We perform experiments using four real-world datasets from different domains: Wikipedia, News, Song lyrics, and European parliament proceedings. We also observe the effect of WEM stability on three downstream tasks: Clustering, POS tagging, and Fairness evaluation. Our experiments indicate that amongst the three WEMs, fastText is the most stable, followed by GloVe and Word2Vec.
format Preprint
id arxiv_https___arxiv_org_abs_2104_08433
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Are Word Embedding Methods Stable and Should We Care About It?
Borah, Angana
Barman, Manash Pratim
Awekar, Amit
Computation and Language
Information Retrieval
A representation learning method is considered stable if it consistently generates similar representation of the given data across multiple runs. Word Embedding Methods (WEMs) are a class of representation learning methods that generate dense vector representation for each word in the given text data. The central idea of this paper is to explore the stability measurement of WEMs using intrinsic evaluation based on word similarity. We experiment with three popular WEMs: Word2Vec, GloVe, and fastText. For stability measurement, we investigate the effect of five parameters involved in training these models. We perform experiments using four real-world datasets from different domains: Wikipedia, News, Song lyrics, and European parliament proceedings. We also observe the effect of WEM stability on three downstream tasks: Clustering, POS tagging, and Fairness evaluation. Our experiments indicate that amongst the three WEMs, fastText is the most stable, followed by GloVe and Word2Vec.
title Are Word Embedding Methods Stable and Should We Care About It?
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2104.08433