Saved in:
Bibliographic Details
Main Authors: Qi, Qianqian, Bagheri, Ayoub, Hessen, David J., van der Heijden, Peter G. M.
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2405.20895
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917326102724608
author Qi, Qianqian
Bagheri, Ayoub
Hessen, David J.
van der Heijden, Peter G. M.
author_facet Qi, Qianqian
Bagheri, Ayoub
Hessen, David J.
van der Heijden, Peter G. M.
contents Popular word embedding methods such as GloVe and Word2Vec are related to the factorization of the pointwise mutual information (PMI) matrix. In this paper, we establish a formal connection between correspondence analysis (CA) and PMI-based word embedding methods. CA is a dimensionality reduction method that uses singular value decomposition (SVD), and we show that CA is mathematically close to the weighted factorization of the PMI matrix. We further introduce variants of CA for word-context matrices, namely CA applied after a square-root transformation (ROOT-CA) and after a fourth-root transformation (ROOTROOT-CA). We analyze the performance of these methods and examine how their success or failure is influenced by extreme values in the decomposed matrix. Although our primary focus is on traditionalstatic word embedding methods, we also include a comparison with a transformer-based encoder (BERT) to situate the results relative to contextual embeddings. Empirical evaluations across multiple corpora and word-similarity benchmarks show that ROOT-CA and ROOTROOT-CA perform slightly better overall than standard PMI-based methods and achieve results competitive with BERT.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20895
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Correspondence Analysis and PMI-Based Word Embeddings: A Comparative Study
Qi, Qianqian
Bagheri, Ayoub
Hessen, David J.
van der Heijden, Peter G. M.
Computation and Language
Popular word embedding methods such as GloVe and Word2Vec are related to the factorization of the pointwise mutual information (PMI) matrix. In this paper, we establish a formal connection between correspondence analysis (CA) and PMI-based word embedding methods. CA is a dimensionality reduction method that uses singular value decomposition (SVD), and we show that CA is mathematically close to the weighted factorization of the PMI matrix. We further introduce variants of CA for word-context matrices, namely CA applied after a square-root transformation (ROOT-CA) and after a fourth-root transformation (ROOTROOT-CA). We analyze the performance of these methods and examine how their success or failure is influenced by extreme values in the decomposed matrix. Although our primary focus is on traditionalstatic word embedding methods, we also include a comparison with a transformer-based encoder (BERT) to situate the results relative to contextual embeddings. Empirical evaluations across multiple corpora and word-similarity benchmarks show that ROOT-CA and ROOTROOT-CA perform slightly better overall than standard PMI-based methods and achieve results competitive with BERT.
title Correspondence Analysis and PMI-Based Word Embeddings: A Comparative Study
topic Computation and Language
url https://arxiv.org/abs/2405.20895