MoSECroT: Model Stitching with Static Word Embeddings for Crosslingual Zero-shot Transfer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Haotian, Liu, Yihong, Ma, Chunlan, Schütze, Hinrich
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911880422883328
author Ye, Haotian
Liu, Yihong
Ma, Chunlan
Schütze, Hinrich
author_facet Ye, Haotian
Liu, Yihong
Ma, Chunlan
Schütze, Hinrich
contents Transformer-based pre-trained language models (PLMs) have achieved remarkable performance in various natural language processing (NLP) tasks. However, pre-training such models can take considerable resources that are almost only available to high-resource languages. On the contrary, static word embeddings are easier to train in terms of computing resources and the amount of data required. In this paper, we introduce MoSECroT Model Stitching with Static Word Embeddings for Crosslingual Zero-shot Transfer), a novel and challenging task that is especially relevant to low-resource languages for which static word embeddings are available. To tackle the task, we present the first framework that leverages relative representations to construct a common space for the embeddings of a source language PLM and the static word embeddings of a target language. In this way, we can train the PLM on source-language training data and perform zero-shot transfer to the target language by simply swapping the embedding layer. However, through extensive experiments on two classification datasets, we show that although our proposed framework is competitive with weak baselines when addressing MoSECroT, it fails to achieve competitive results compared with some strong baselines. In this paper, we attempt to explain this negative result and provide several thoughts on possible improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2401_04821
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MoSECroT: Model Stitching with Static Word Embeddings for Crosslingual Zero-shot Transfer
Ye, Haotian
Liu, Yihong
Ma, Chunlan
Schütze, Hinrich
Computation and Language
Artificial Intelligence
Transformer-based pre-trained language models (PLMs) have achieved remarkable performance in various natural language processing (NLP) tasks. However, pre-training such models can take considerable resources that are almost only available to high-resource languages. On the contrary, static word embeddings are easier to train in terms of computing resources and the amount of data required. In this paper, we introduce MoSECroT Model Stitching with Static Word Embeddings for Crosslingual Zero-shot Transfer), a novel and challenging task that is especially relevant to low-resource languages for which static word embeddings are available. To tackle the task, we present the first framework that leverages relative representations to construct a common space for the embeddings of a source language PLM and the static word embeddings of a target language. In this way, we can train the PLM on source-language training data and perform zero-shot transfer to the target language by simply swapping the embedding layer. However, through extensive experiments on two classification datasets, we show that although our proposed framework is competitive with weak baselines when addressing MoSECroT, it fails to achieve competitive results compared with some strong baselines. In this paper, we attempt to explain this negative result and provide several thoughts on possible improvement.
title MoSECroT: Model Stitching with Static Word Embeddings for Crosslingual Zero-shot Transfer
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.04821