| _version_ | 1866902336057638912 |
|---|---|
| author | McGill, Euan Chiruzzo, Luis Saggion, Horacio |
| author_facet | McGill, Euan Chiruzzo, Luis Saggion, Horacio |
| contents | <p>Word2Vec-based pre-trained word or gloss embedding models for Sign Languages. These are created by modifying the weights of pre-existing spoken language word embedding models to their SL counterparts, owing to their shared lexemes in Sign Language gloss. A full description of how they are created may be found in the companion paper (McGill et al., 2024).</p> <p>Languages included:</p> <ul> <li>American Sign Language (ASL) <ul> <li>3M embeddings, 5.1k unique signs, 300D</li> </ul> </li> <li>Spanish Sign Language (LSE) <ul> <li>1M embeddings, 1.2k unique signs, 300D</li> </ul> </li> <li>Sign Language of the Netherlands (NGT) <ul> <li>627k embeddings, 4.1k unique signs, 320D</li> </ul> </li> <li>Finnish Sign Language (FinSL) <ul> <li>247k embeddings, 3.1k unique signs, 100D</li> </ul> </li> </ul> <p>All four models are in .txt format, with the binary versions available for ASL and LSE. Don't hesitate to contact for further information.</p> <p> </p> <p>Abstract from original paper:</p> <p>This paper explores a novel method to modify existing pre-trained word embedding models of spoken languages for Sign Language glosses. These newly-generated embeddings are described, visualised, and then used in the encoder and/or decoder of models for the Text2Gloss and Gloss2Text task of machine translation. In two translation settings (one including data augmentation-based pre-training and a baseline), we find that bootstrapped word embeddings for glosses improve translation across four Signed/spoken language pairs. Many improvements are statistically significant, including those where the bootstrapped gloss embedding models are used.</p> <p>Languages included: American Sign Language, Finnish Sign Language, Spanish Sign Language, Sign Language of The Netherlands.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_15344009 |
| institution | Zenodo |
| language | |
| publishDate | 2024 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Sign Language Gloss Embedding Models McGill, Euan Chiruzzo, Luis Saggion, Horacio word vectors pre-trained word embeddings lexical semantics sign language Sign Language Sign language machine translation <p>Word2Vec-based pre-trained word or gloss embedding models for Sign Languages. These are created by modifying the weights of pre-existing spoken language word embedding models to their SL counterparts, owing to their shared lexemes in Sign Language gloss. A full description of how they are created may be found in the companion paper (McGill et al., 2024).</p> <p>Languages included:</p> <ul> <li>American Sign Language (ASL) <ul> <li>3M embeddings, 5.1k unique signs, 300D</li> </ul> </li> <li>Spanish Sign Language (LSE) <ul> <li>1M embeddings, 1.2k unique signs, 300D</li> </ul> </li> <li>Sign Language of the Netherlands (NGT) <ul> <li>627k embeddings, 4.1k unique signs, 320D</li> </ul> </li> <li>Finnish Sign Language (FinSL) <ul> <li>247k embeddings, 3.1k unique signs, 100D</li> </ul> </li> </ul> <p>All four models are in .txt format, with the binary versions available for ASL and LSE. Don't hesitate to contact for further information.</p> <p> </p> <p>Abstract from original paper:</p> <p>This paper explores a novel method to modify existing pre-trained word embedding models of spoken languages for Sign Language glosses. These newly-generated embeddings are described, visualised, and then used in the encoder and/or decoder of models for the Text2Gloss and Gloss2Text task of machine translation. In two translation settings (one including data augmentation-based pre-training and a baseline), we find that bootstrapped word embeddings for glosses improve translation across four Signed/spoken language pairs. Many improvements are statistically significant, including those where the bootstrapped gloss embedding models are used.</p> <p>Languages included: American Sign Language, Finnish Sign Language, Spanish Sign Language, Sign Language of The Netherlands.</p> |
| title | Sign Language Gloss Embedding Models |
| topic | word vectors pre-trained word embeddings lexical semantics sign language Sign Language Sign language machine translation |
| url | https://doi.org/10.5281/zenodo.15344009 |