Unsupervised Bilingual Lexicon Induction for Low Resource Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rathnayake, Charitha, Thilakarathna, P. R. S., Nethmini, Uthpala, Kaur, Rishemjith, Ranathunga, Surangika
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929644472631296
author Rathnayake, Charitha
Thilakarathna, P. R. S.
Nethmini, Uthpala
Kaur, Rishemjith
Ranathunga, Surangika
author_facet Rathnayake, Charitha
Thilakarathna, P. R. S.
Nethmini, Uthpala
Kaur, Rishemjith
Ranathunga, Surangika
contents Bilingual lexicons play a crucial role in various Natural Language Processing tasks. However, many low-resource languages (LRLs) do not have such lexicons, and due to the same reason, cannot benefit from the supervised Bilingual Lexicon Induction (BLI) techniques. To address this, unsupervised BLI (UBLI) techniques were introduced. A prominent technique in this line is structure-based UBLI. It is an iterative method, where a seed lexicon, which is initially learned from monolingual embeddings is iteratively improved. There have been numerous improvements to this core idea, however they have been experimented with independently of each other. In this paper, we investigate whether using these techniques simultaneously would lead to equal gains. We use the unsupervised version of VecMap, a commonly used structure-based UBLI framework, and carry out a comprehensive set of experiments using the LRL pairs, English-Sinhala, English-Tamil, and English-Punjabi. These experiments helped us to identify the best combination of the extensions. We also release bilingual dictionaries for English-Sinhala and English-Punjabi.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16894
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unsupervised Bilingual Lexicon Induction for Low Resource Languages
Rathnayake, Charitha
Thilakarathna, P. R. S.
Nethmini, Uthpala
Kaur, Rishemjith
Ranathunga, Surangika
Computation and Language
Bilingual lexicons play a crucial role in various Natural Language Processing tasks. However, many low-resource languages (LRLs) do not have such lexicons, and due to the same reason, cannot benefit from the supervised Bilingual Lexicon Induction (BLI) techniques. To address this, unsupervised BLI (UBLI) techniques were introduced. A prominent technique in this line is structure-based UBLI. It is an iterative method, where a seed lexicon, which is initially learned from monolingual embeddings is iteratively improved. There have been numerous improvements to this core idea, however they have been experimented with independently of each other. In this paper, we investigate whether using these techniques simultaneously would lead to equal gains. We use the unsupervised version of VecMap, a commonly used structure-based UBLI framework, and carry out a comprehensive set of experiments using the LRL pairs, English-Sinhala, English-Tamil, and English-Punjabi. These experiments helped us to identify the best combination of the extensions. We also release bilingual dictionaries for English-Sinhala and English-Punjabi.
title Unsupervised Bilingual Lexicon Induction for Low Resource Languages
topic Computation and Language
url https://arxiv.org/abs/2412.16894