Saved in:
| Main Author: | |
|---|---|
| Format: | Recurso digital |
| Language: | |
| Published: |
Zenodo
2026
|
| Online Access: | https://doi.org/10.5281/zenodo.18258689 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866901915246264320 |
|---|---|
| author | Journal of Theoretical and Applied Information Technology |
| author_facet | Journal of Theoretical and Applied Information Technology |
| contents | <p><span>Coreference resolution which establishes whether different kind of expressions in a text relate to the same thing, is an essential part of natural language understanding. Although this task has seen substantial development in high-resource languages, particularly through the use of large annotated datasets and pretrained models development. Languages like Tamil under explored because of the limited annotated datasets and specialized language models. To overcome this, a zero-shot method for coreference resolution in Tamil by (MuRIL,) a multilingual pretrained language model is proposed.The framework adapts span-based neural architectures, using contextual span representations, cosine similarity-based scoring, and unsupervised agglomerative clustering to detect and link coreferent mentions efficiently.To facilitate evaluation, a new manually annotated dataset, TGAP (Tamil GAP-style Coreference Dataset) was developed drawn from diverse text sources. Experimental evaluation compared MuRIL with other multilingual baselines using standard metrics such as MUC, B³, CEAF, and LEA F1 scores. Results shows that MuRIL gives the highest overall F1 score (0.67), performing well with other models underscoring the value of language-specific pretraining. The findings shows that a zero-shot, annotation-free framework can achieve competitive performance for Tamil coreference resolution.</span></p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18258689 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | ADDRESSING THE DATA SCARCITY OF TAMIL: A ZERO-SHOT COREFERENCE RESOLUTION APPROACH WITH MURIL Journal of Theoretical and Applied Information Technology <p><span>Coreference resolution which establishes whether different kind of expressions in a text relate to the same thing, is an essential part of natural language understanding. Although this task has seen substantial development in high-resource languages, particularly through the use of large annotated datasets and pretrained models development. Languages like Tamil under explored because of the limited annotated datasets and specialized language models. To overcome this, a zero-shot method for coreference resolution in Tamil by (MuRIL,) a multilingual pretrained language model is proposed.The framework adapts span-based neural architectures, using contextual span representations, cosine similarity-based scoring, and unsupervised agglomerative clustering to detect and link coreferent mentions efficiently.To facilitate evaluation, a new manually annotated dataset, TGAP (Tamil GAP-style Coreference Dataset) was developed drawn from diverse text sources. Experimental evaluation compared MuRIL with other multilingual baselines using standard metrics such as MUC, B³, CEAF, and LEA F1 scores. Results shows that MuRIL gives the highest overall F1 score (0.67), performing well with other models underscoring the value of language-specific pretraining. The findings shows that a zero-shot, annotation-free framework can achieve competitive performance for Tamil coreference resolution.</span></p> |
| title | ADDRESSING THE DATA SCARCITY OF TAMIL: A ZERO-SHOT COREFERENCE RESOLUTION APPROACH WITH MURIL |
| url | https://doi.org/10.5281/zenodo.18258689 |