SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Goworek, Roksana, Karlcut, Harpal, Shezad, Muhammad, Darshana, Nijaguna, Mane, Abhishek, Bondada, Syam, Sikka, Raghav, Mammadov, Ulvi, Allahverdiyev, Rauf, Purighella, Sriram, Gupta, Paridhi, Ndegwa, Muhinyia, Dubossarsky, Haim
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909699618635776
author Goworek, Roksana
Karlcut, Harpal
Shezad, Muhammad
Darshana, Nijaguna
Mane, Abhishek
Bondada, Syam
Sikka, Raghav
Mammadov, Ulvi
Allahverdiyev, Rauf
Purighella, Sriram
Gupta, Paridhi
Ndegwa, Muhinyia
Dubossarsky, Haim
author_facet Goworek, Roksana
Karlcut, Harpal
Shezad, Muhammad
Darshana, Nijaguna
Mane, Abhishek
Bondada, Syam
Sikka, Raghav
Mammadov, Ulvi
Allahverdiyev, Rauf
Purighella, Sriram
Gupta, Paridhi
Ndegwa, Muhinyia
Dubossarsky, Haim
contents This paper addresses the critical need for high-quality evaluation datasets in low-resource languages to advance cross-lingual transfer. While cross-lingual transfer offers a key strategy for leveraging multilingual pretraining to expand language technologies to understudied and typologically diverse languages, its effectiveness is dependent on quality and suitable benchmarks. We release new sense-annotated datasets of sentences containing polysemous words, spanning ten low-resource languages across diverse language families and scripts. To facilitate dataset creation, the paper presents a demonstrably beneficial semi-automatic annotation method. The utility of the datasets is demonstrated through Word-in-Context (WiC) formatted experiments that evaluate transfer on these low-resource languages. Results highlight the importance of targeted dataset creation and evaluation for effective polysemy disambiguation in low-resource settings and transfer studies. The released datasets and code aim to support further research into fair, robust, and truly multilingual NLP.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23714
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods
Goworek, Roksana
Karlcut, Harpal
Shezad, Muhammad
Darshana, Nijaguna
Mane, Abhishek
Bondada, Syam
Sikka, Raghav
Mammadov, Ulvi
Allahverdiyev, Rauf
Purighella, Sriram
Gupta, Paridhi
Ndegwa, Muhinyia
Dubossarsky, Haim
Computation and Language
Artificial Intelligence
This paper addresses the critical need for high-quality evaluation datasets in low-resource languages to advance cross-lingual transfer. While cross-lingual transfer offers a key strategy for leveraging multilingual pretraining to expand language technologies to understudied and typologically diverse languages, its effectiveness is dependent on quality and suitable benchmarks. We release new sense-annotated datasets of sentences containing polysemous words, spanning ten low-resource languages across diverse language families and scripts. To facilitate dataset creation, the paper presents a demonstrably beneficial semi-automatic annotation method. The utility of the datasets is demonstrated through Word-in-Context (WiC) formatted experiments that evaluate transfer on these low-resource languages. Results highlight the importance of targeted dataset creation and evaluation for effective polysemy disambiguation in low-resource settings and transfer studies. The released datasets and code aim to support further research into fair, robust, and truly multilingual NLP.
title SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.23714