EMS: Efficient and Effective Massively Multilingual Sentence Embedding Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Zhuoyuan, Chu, Chenhui, Kurohashi, Sadao
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911892842217472
author Mao, Zhuoyuan
Chu, Chenhui
Kurohashi, Sadao
author_facet Mao, Zhuoyuan
Chu, Chenhui
Kurohashi, Sadao
contents Massively multilingual sentence representation models, e.g., LASER, SBERT-distill, and LaBSE, help significantly improve cross-lingual downstream tasks. However, the use of a large amount of data or inefficient model architectures results in heavy computation to train a new model according to our preferred languages and domains. To resolve this issue, we introduce efficient and effective massively multilingual sentence embedding (EMS), using cross-lingual token-level reconstruction (XTR) and sentence-level contrastive learning as training objectives. Compared with related studies, the proposed model can be efficiently trained using significantly fewer parallel sentences and GPU computation resources. Empirical results showed that the proposed model significantly yields better or comparable results with regard to cross-lingual sentence retrieval, zero-shot cross-lingual genre classification, and sentiment classification. Ablative analyses demonstrated the efficiency and effectiveness of each component of the proposed model. We release the codes for model training and the EMS pre-trained sentence embedding model, which supports 62 languages ( https://github.com/Mao-KU/EMS ).
format Preprint
id arxiv_https___arxiv_org_abs_2205_15744
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle EMS: Efficient and Effective Massively Multilingual Sentence Embedding Learning
Mao, Zhuoyuan
Chu, Chenhui
Kurohashi, Sadao
Computation and Language
Massively multilingual sentence representation models, e.g., LASER, SBERT-distill, and LaBSE, help significantly improve cross-lingual downstream tasks. However, the use of a large amount of data or inefficient model architectures results in heavy computation to train a new model according to our preferred languages and domains. To resolve this issue, we introduce efficient and effective massively multilingual sentence embedding (EMS), using cross-lingual token-level reconstruction (XTR) and sentence-level contrastive learning as training objectives. Compared with related studies, the proposed model can be efficiently trained using significantly fewer parallel sentences and GPU computation resources. Empirical results showed that the proposed model significantly yields better or comparable results with regard to cross-lingual sentence retrieval, zero-shot cross-lingual genre classification, and sentiment classification. Ablative analyses demonstrated the efficiency and effectiveness of each component of the proposed model. We release the codes for model training and the EMS pre-trained sentence embedding model, which supports 62 languages ( https://github.com/Mao-KU/EMS ).
title EMS: Efficient and Effective Massively Multilingual Sentence Embedding Learning
topic Computation and Language
url https://arxiv.org/abs/2205.15744