Salvato in:
Dettagli Bibliografici
Autori principali: Park, Minsu, Choi, Seyeon, Choi, Chanyeol, Kim, Jun-Seong, Sohn, Jy-yong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:https://arxiv.org/abs/2405.16155
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914814151884800
author Park, Minsu
Choi, Seyeon
Choi, Chanyeol
Kim, Jun-Seong
Sohn, Jy-yong
author_facet Park, Minsu
Choi, Seyeon
Choi, Chanyeol
Kim, Jun-Seong
Sohn, Jy-yong
contents Making decent multi-lingual sentence representations is critical to achieve high performances in cross-lingual downstream tasks. In this work, we propose a novel method to align multi-lingual embeddings based on the similarity of sentences measured by a pre-trained mono-lingual embedding model. Given translation sentence pairs, we train a multi-lingual model in a way that the similarity between cross-lingual embeddings follows the similarity of sentences measured at the mono-lingual teacher model. Our method can be considered as contrastive learning with soft labels defined as the similarity between sentences. Our experimental results on five languages show that our contrastive loss with soft labels far outperforms conventional contrastive loss with hard labels in various benchmarks for bitext mining tasks and STS tasks. In addition, our method outperforms existing multi-lingual embeddings including LaBSE, for Tatoeba dataset. The code is available at https://github.com/YAI12xLinq-B/IMASCL
format Preprint
id arxiv_https___arxiv_org_abs_2405_16155
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Multi-lingual Alignment Through Soft Contrastive Learning
Park, Minsu
Choi, Seyeon
Choi, Chanyeol
Kim, Jun-Seong
Sohn, Jy-yong
Computation and Language
Making decent multi-lingual sentence representations is critical to achieve high performances in cross-lingual downstream tasks. In this work, we propose a novel method to align multi-lingual embeddings based on the similarity of sentences measured by a pre-trained mono-lingual embedding model. Given translation sentence pairs, we train a multi-lingual model in a way that the similarity between cross-lingual embeddings follows the similarity of sentences measured at the mono-lingual teacher model. Our method can be considered as contrastive learning with soft labels defined as the similarity between sentences. Our experimental results on five languages show that our contrastive loss with soft labels far outperforms conventional contrastive loss with hard labels in various benchmarks for bitext mining tasks and STS tasks. In addition, our method outperforms existing multi-lingual embeddings including LaBSE, for Tatoeba dataset. The code is available at https://github.com/YAI12xLinq-B/IMASCL
title Improving Multi-lingual Alignment Through Soft Contrastive Learning
topic Computation and Language
url https://arxiv.org/abs/2405.16155