Saved in:
Bibliographic Details
Main Author: Journal of Theoretical and Applied Information Technology
Format: Recurso digital
Language:
Published: Zenodo 2026
Online Access:https://doi.org/10.5281/zenodo.18258689
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866901915246264320
author Journal of Theoretical and Applied Information Technology
author_facet Journal of Theoretical and Applied Information Technology
contents <p><span>Coreference resolution which establishes whether different kind of expressions in a text relate to the same thing, is an essential part of natural language understanding. Although this task has seen substantial development in high-resource languages, particularly through the use of large annotated datasets and pretrained models development. Languages like Tamil under explored because of the limited annotated datasets and specialized language models. To overcome this, a zero-shot method for coreference resolution in Tamil by (MuRIL,) a multilingual pretrained language model is proposed.The framework adapts span-based neural architectures, using contextual span representations, cosine similarity-based scoring, and unsupervised agglomerative clustering to detect and link coreferent mentions efficiently.To facilitate evaluation, a new manually annotated dataset, TGAP (Tamil GAP-style Coreference Dataset) was developed drawn from diverse text sources. Experimental evaluation compared MuRIL with other multilingual baselines using standard metrics such as MUC, B³, CEAF, and LEA F1 scores. Results shows that MuRIL gives the highest overall F1 score (0.67), performing well with other models underscoring the value of language-specific pretraining. The findings shows that a zero-shot, annotation-free framework can achieve competitive performance for Tamil coreference resolution.</span></p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18258689
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle ADDRESSING THE DATA SCARCITY OF TAMIL: A ZERO-SHOT COREFERENCE RESOLUTION APPROACH WITH MURIL
Journal of Theoretical and Applied Information Technology
<p><span>Coreference resolution which establishes whether different kind of expressions in a text relate to the same thing, is an essential part of natural language understanding. Although this task has seen substantial development in high-resource languages, particularly through the use of large annotated datasets and pretrained models development. Languages like Tamil under explored because of the limited annotated datasets and specialized language models. To overcome this, a zero-shot method for coreference resolution in Tamil by (MuRIL,) a multilingual pretrained language model is proposed.The framework adapts span-based neural architectures, using contextual span representations, cosine similarity-based scoring, and unsupervised agglomerative clustering to detect and link coreferent mentions efficiently.To facilitate evaluation, a new manually annotated dataset, TGAP (Tamil GAP-style Coreference Dataset) was developed drawn from diverse text sources. Experimental evaluation compared MuRIL with other multilingual baselines using standard metrics such as MUC, B³, CEAF, and LEA F1 scores. Results shows that MuRIL gives the highest overall F1 score (0.67), performing well with other models underscoring the value of language-specific pretraining. The findings shows that a zero-shot, annotation-free framework can achieve competitive performance for Tamil coreference resolution.</span></p>
title ADDRESSING THE DATA SCARCITY OF TAMIL: A ZERO-SHOT COREFERENCE RESOLUTION APPROACH WITH MURIL
url https://doi.org/10.5281/zenodo.18258689