Lithuanian word2vec embeddings trained on OpenSubtitles

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Authors: Grim, Philip, Buchanan, Erin
Format: Recurso digital
Published: Zenodo 2025
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866901610048782336
author Grim, Philip
Buchanan, Erin
author_facet Grim, Philip
Buchanan, Erin
contents <p>This dataset contains the subs2vec embeddings for Lithuanian, as presented in <a href="https://zenodo.org/records/17243814" target="_blank" rel="noopener">https://zenodo.org/records/17243814</a>. The embeddings were trained on large-scale subtitle corpora and represent semantic vector spaces derived from naturalistic language use in films and television from the OpenSubtitles 2018 datasets: <a href="https://opus.nlpl.eu/OpenSubtitles/corpus/version/OpenSubtitles" target="_blank" rel="noopener">https://opus.nlpl.eu/OpenSubtitles/corpus/version/OpenSubtitles</a>. </p> <p>For this language, we provide all embedding variants explored in the study. Specifically, the dataset includes vectors generated under different combinations of:</p> <ul> <li>Dimensionality: multiple vector sizes (e.g., 100, 200, 300, …)</li> <li>Window size: varying context windows (e.g., 2, 5, 10, …)</li> <li>Each file corresponds to a unique configuration (dimension × window size). </li> </ul> <p>Each file contains the vocabulary for that language (column 1) and then the embedding values (columns 2 through dimension size + 1). </p> <p>If you use this dataset, please cite:</p> <ul> <li>Manuscript: <a href="https://doi.org/10.5281/zenodo.17243812" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.17243812</a> </li> <li>Data: This Zenodo dataset (using the DOI provided here)</li> </ul>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_17559525
institution Zenodo
language
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle Lithuanian word2vec embeddings trained on OpenSubtitles
Grim, Philip
Buchanan, Erin
<p>This dataset contains the subs2vec embeddings for Lithuanian, as presented in <a href="https://zenodo.org/records/17243814" target="_blank" rel="noopener">https://zenodo.org/records/17243814</a>. The embeddings were trained on large-scale subtitle corpora and represent semantic vector spaces derived from naturalistic language use in films and television from the OpenSubtitles 2018 datasets: <a href="https://opus.nlpl.eu/OpenSubtitles/corpus/version/OpenSubtitles" target="_blank" rel="noopener">https://opus.nlpl.eu/OpenSubtitles/corpus/version/OpenSubtitles</a>. </p> <p>For this language, we provide all embedding variants explored in the study. Specifically, the dataset includes vectors generated under different combinations of:</p> <ul> <li>Dimensionality: multiple vector sizes (e.g., 100, 200, 300, …)</li> <li>Window size: varying context windows (e.g., 2, 5, 10, …)</li> <li>Each file corresponds to a unique configuration (dimension × window size). </li> </ul> <p>Each file contains the vocabulary for that language (column 1) and then the embedding values (columns 2 through dimension size + 1). </p> <p>If you use this dataset, please cite:</p> <ul> <li>Manuscript: <a href="https://doi.org/10.5281/zenodo.17243812" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.17243812</a> </li> <li>Data: This Zenodo dataset (using the DOI provided here)</li> </ul>
title Lithuanian word2vec embeddings trained on OpenSubtitles
url https://doi.org/10.5281/zenodo.17559525