| _version_ | 1866901610048782336 |
|---|---|
| author | Grim, Philip Buchanan, Erin |
| author_facet | Grim, Philip Buchanan, Erin |
| contents | <p>This dataset contains the subs2vec embeddings for Lithuanian, as presented in <a href="https://zenodo.org/records/17243814" target="_blank" rel="noopener">https://zenodo.org/records/17243814</a>. The embeddings were trained on large-scale subtitle corpora and represent semantic vector spaces derived from naturalistic language use in films and television from the OpenSubtitles 2018 datasets: <a href="https://opus.nlpl.eu/OpenSubtitles/corpus/version/OpenSubtitles" target="_blank" rel="noopener">https://opus.nlpl.eu/OpenSubtitles/corpus/version/OpenSubtitles</a>. </p> <p>For this language, we provide all embedding variants explored in the study. Specifically, the dataset includes vectors generated under different combinations of:</p> <ul> <li>Dimensionality: multiple vector sizes (e.g., 100, 200, 300, …)</li> <li>Window size: varying context windows (e.g., 2, 5, 10, …)</li> <li>Each file corresponds to a unique configuration (dimension × window size). </li> </ul> <p>Each file contains the vocabulary for that language (column 1) and then the embedding values (columns 2 through dimension size + 1). </p> <p>If you use this dataset, please cite:</p> <ul> <li>Manuscript: <a href="https://doi.org/10.5281/zenodo.17243812" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.17243812</a> </li> <li>Data: This Zenodo dataset (using the DOI provided here)</li> </ul> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_17559525 |
| institution | Zenodo |
| language | |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Lithuanian word2vec embeddings trained on OpenSubtitles Grim, Philip Buchanan, Erin <p>This dataset contains the subs2vec embeddings for Lithuanian, as presented in <a href="https://zenodo.org/records/17243814" target="_blank" rel="noopener">https://zenodo.org/records/17243814</a>. The embeddings were trained on large-scale subtitle corpora and represent semantic vector spaces derived from naturalistic language use in films and television from the OpenSubtitles 2018 datasets: <a href="https://opus.nlpl.eu/OpenSubtitles/corpus/version/OpenSubtitles" target="_blank" rel="noopener">https://opus.nlpl.eu/OpenSubtitles/corpus/version/OpenSubtitles</a>. </p> <p>For this language, we provide all embedding variants explored in the study. Specifically, the dataset includes vectors generated under different combinations of:</p> <ul> <li>Dimensionality: multiple vector sizes (e.g., 100, 200, 300, …)</li> <li>Window size: varying context windows (e.g., 2, 5, 10, …)</li> <li>Each file corresponds to a unique configuration (dimension × window size). </li> </ul> <p>Each file contains the vocabulary for that language (column 1) and then the embedding values (columns 2 through dimension size + 1). </p> <p>If you use this dataset, please cite:</p> <ul> <li>Manuscript: <a href="https://doi.org/10.5281/zenodo.17243812" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.17243812</a> </li> <li>Data: This Zenodo dataset (using the DOI provided here)</li> </ul> |
| title | Lithuanian word2vec embeddings trained on OpenSubtitles |
| url | https://doi.org/10.5281/zenodo.17559525 |