Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.02056 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917892837081088 |
|---|---|
| author | Ranathunga, Surangika Ranasinghea, Asanka Shamala, Janaka Dandeniyaa, Ayodya Galappaththia, Rashmi Samaraweeraa, Malithi |
| author_facet | Ranathunga, Surangika Ranasinghea, Asanka Shamala, Janaka Dandeniyaa, Ayodya Galappaththia, Rashmi Samaraweeraa, Malithi |
| contents | This paper presents a multi-way parallel English-Tamil-Sinhala corpus annotated with Named Entities (NEs), where Sinhala and Tamil are low-resource languages. Using pre-trained multilingual Language Models (mLMs), we establish new benchmark Named Entity Recognition (NER) results on this dataset for Sinhala and Tamil. We also carry out a detailed investigation on the NER capabilities of different types of mLMs. Finally, we demonstrate the utility of our NER system on a low-resource Neural Machine Translation (NMT) task. Our dataset is publicly released: https://github.com/suralk/multiNER. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_02056 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala Ranathunga, Surangika Ranasinghea, Asanka Shamala, Janaka Dandeniyaa, Ayodya Galappaththia, Rashmi Samaraweeraa, Malithi Computation and Language This paper presents a multi-way parallel English-Tamil-Sinhala corpus annotated with Named Entities (NEs), where Sinhala and Tamil are low-resource languages. Using pre-trained multilingual Language Models (mLMs), we establish new benchmark Named Entity Recognition (NER) results on this dataset for Sinhala and Tamil. We also carry out a detailed investigation on the NER capabilities of different types of mLMs. Finally, we demonstrate the utility of our NER system on a low-resource Neural Machine Translation (NMT) task. Our dataset is publicly released: https://github.com/suralk/multiNER. |
| title | A Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2412.02056 |