Spoken Word2Vec: Learning Skipgram Embeddings from Speech

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sayeed, Mohammad Amaan, Aldarmaki, Hanan
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914855311638528
author Sayeed, Mohammad Amaan
Aldarmaki, Hanan
author_facet Sayeed, Mohammad Amaan
Aldarmaki, Hanan
contents Text word embeddings that encode distributional semantics work by modeling contextual similarities of frequently occurring words. Acoustic word embeddings, on the other hand, typically encode low-level phonetic similarities. Semantic embeddings for spoken words have been previously explored using analogous algorithms to Word2Vec, but the resulting vectors still mainly encoded phonetic rather than semantic features. In this paper, we examine the assumptions and architectures used in previous works and show experimentally how shallow skipgram-like algorithms fail to encode distributional semantics when the input units are acoustically correlated. We illustrate the potential of an alternative deep end-to-end variant of the model and examine the effects on the resulting embeddings, showing positive results of semantic relatedness in the embedding space.
format Preprint
id arxiv_https___arxiv_org_abs_2311_09319
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Spoken Word2Vec: Learning Skipgram Embeddings from Speech
Sayeed, Mohammad Amaan
Aldarmaki, Hanan
Computation and Language
Artificial Intelligence
Text word embeddings that encode distributional semantics work by modeling contextual similarities of frequently occurring words. Acoustic word embeddings, on the other hand, typically encode low-level phonetic similarities. Semantic embeddings for spoken words have been previously explored using analogous algorithms to Word2Vec, but the resulting vectors still mainly encoded phonetic rather than semantic features. In this paper, we examine the assumptions and architectures used in previous works and show experimentally how shallow skipgram-like algorithms fail to encode distributional semantics when the input units are acoustically correlated. We illustrate the potential of an alternative deep end-to-end variant of the model and examine the effects on the resulting embeddings, showing positive results of semantic relatedness in the embedding space.
title Spoken Word2Vec: Learning Skipgram Embeddings from Speech
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2311.09319