Homonym Sense Disambiguation in the Georgian Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Melikidze, Davit, Gamkrelidze, Alexander
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916231818248192
author Melikidze, Davit
Gamkrelidze, Alexander
author_facet Melikidze, Davit
Gamkrelidze, Alexander
contents This research proposes a novel approach to the Word Sense Disambiguation (WSD) task in the Georgian language, based on supervised fine-tuning of a pre-trained Large Language Model (LLM) on a dataset formed by filtering the Georgian Common Crawls corpus. The dataset is used to train a classifier for words with multiple senses. Additionally, we present experimental results of using LSTM for WSD. Accurately disambiguating homonyms is crucial in natural language processing. Georgian, an agglutinative language belonging to the Kartvelian language family, presents unique challenges in this context. The aim of this paper is to highlight the specific problems concerning homonym disambiguation in the Georgian language and to present our approach to solving them. The techniques discussed in the article achieve 95% accuracy for predicting lexical meanings of homonyms using a hand-classified dataset of over 7500 sentences.
format Preprint
id arxiv_https___arxiv_org_abs_2405_00710
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Homonym Sense Disambiguation in the Georgian Language
Melikidze, Davit
Gamkrelidze, Alexander
Computation and Language
Machine Learning
This research proposes a novel approach to the Word Sense Disambiguation (WSD) task in the Georgian language, based on supervised fine-tuning of a pre-trained Large Language Model (LLM) on a dataset formed by filtering the Georgian Common Crawls corpus. The dataset is used to train a classifier for words with multiple senses. Additionally, we present experimental results of using LSTM for WSD. Accurately disambiguating homonyms is crucial in natural language processing. Georgian, an agglutinative language belonging to the Kartvelian language family, presents unique challenges in this context. The aim of this paper is to highlight the specific problems concerning homonym disambiguation in the Georgian language and to present our approach to solving them. The techniques discussed in the article achieve 95% accuracy for predicting lexical meanings of homonyms using a hand-classified dataset of over 7500 sentences.
title Homonym Sense Disambiguation in the Georgian Language
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2405.00710