LexGen: Domain-aware Multilingual Lexicon Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maheshwari, Ayush, Singh, Atul Kumar, NJ, Karthika, Bhatt, Krishnakant, Jyothi, Preethi, Ramakrishnan, Ganesh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909630673715200
author Maheshwari, Ayush
Singh, Atul Kumar
NJ, Karthika
Bhatt, Krishnakant
Jyothi, Preethi
Ramakrishnan, Ganesh
author_facet Maheshwari, Ayush
Singh, Atul Kumar
NJ, Karthika
Bhatt, Krishnakant
Jyothi, Preethi
Ramakrishnan, Ganesh
contents Lexicon or dictionary generation across domains has the potential for societal impact, as it can potentially enhance information accessibility for a diverse user base while preserving language identity. Prior work in the field primarily focuses on bilingual lexical induction, which deals with word alignments using mapping or corpora-based approaches. However, these approaches do not cater to domain-specific lexicon generation that consists of domain-specific terminology. This task becomes particularly important in specialized medical, engineering, and other technical domains, owing to the highly infrequent usage of the terms and scarcity of data involving domain-specific terms especially for low/mid-resource languages. In this paper, we propose a new model to generate dictionary words for $6$ Indian languages in the multi-domain setting. Our model consists of domain-specific and domain-generic layers that encode information, and these layers are invoked via a learnable routing technique. We also release a new benchmark dataset consisting of >75K translation pairs across 6 Indian languages spanning 8 diverse domains.We conduct both zero-shot and few-shot experiments across multiple domains to show the efficacy of our proposed model in generalizing to unseen domains and unseen languages. Additionally, we also perform a post-hoc human evaluation on unseen languages. The source code and dataset is present at https://github.com/Atulkmrsingh/lexgen.
format Preprint
id arxiv_https___arxiv_org_abs_2405_11200
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LexGen: Domain-aware Multilingual Lexicon Generation
Maheshwari, Ayush
Singh, Atul Kumar
NJ, Karthika
Bhatt, Krishnakant
Jyothi, Preethi
Ramakrishnan, Ganesh
Computation and Language
Lexicon or dictionary generation across domains has the potential for societal impact, as it can potentially enhance information accessibility for a diverse user base while preserving language identity. Prior work in the field primarily focuses on bilingual lexical induction, which deals with word alignments using mapping or corpora-based approaches. However, these approaches do not cater to domain-specific lexicon generation that consists of domain-specific terminology. This task becomes particularly important in specialized medical, engineering, and other technical domains, owing to the highly infrequent usage of the terms and scarcity of data involving domain-specific terms especially for low/mid-resource languages. In this paper, we propose a new model to generate dictionary words for $6$ Indian languages in the multi-domain setting. Our model consists of domain-specific and domain-generic layers that encode information, and these layers are invoked via a learnable routing technique. We also release a new benchmark dataset consisting of >75K translation pairs across 6 Indian languages spanning 8 diverse domains.We conduct both zero-shot and few-shot experiments across multiple domains to show the efficacy of our proposed model in generalizing to unseen domains and unseen languages. Additionally, we also perform a post-hoc human evaluation on unseen languages. The source code and dataset is present at https://github.com/Atulkmrsingh/lexgen.
title LexGen: Domain-aware Multilingual Lexicon Generation
topic Computation and Language
url https://arxiv.org/abs/2405.11200