Neural Compound-Word (Sandhi) Generation and Splitting in Sanskrit Language

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dave, Sushant, Singh, Arun Kumar, P., Prathosh A., Lall, Brejesh
Natura: Preprint
Pubblicazione: 2020
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910587187888128
author Dave, Sushant
Singh, Arun Kumar
P., Prathosh A.
Lall, Brejesh
author_facet Dave, Sushant
Singh, Arun Kumar
P., Prathosh A.
Lall, Brejesh
contents This paper describes neural network based approaches to the process of the formation and splitting of word-compounding, respectively known as the Sandhi and Vichchhed, in Sanskrit language. Sandhi is an important idea essential to morphological analysis of Sanskrit texts. Sandhi leads to word transformations at word boundaries. The rules of Sandhi formation are well defined but complex, sometimes optional and in some cases, require knowledge about the nature of the words being compounded. Sandhi split or Vichchhed is an even more difficult task given its non uniqueness and context dependence. In this work, we propose the route of formulating the problem as a sequence to sequence prediction task, using modern deep learning techniques. Being the first fully data driven technique, we demonstrate that our model has an accuracy better than the existing methods on multiple standard datasets, despite not using any additional lexical or morphological resources. The code is being made available at https://github.com/IITD-DataScience/Sandhi_Prakarana
format Preprint
id arxiv_https___arxiv_org_abs_2010_12940
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Neural Compound-Word (Sandhi) Generation and Splitting in Sanskrit Language
Dave, Sushant
Singh, Arun Kumar
P., Prathosh A.
Lall, Brejesh
Computation and Language
This paper describes neural network based approaches to the process of the formation and splitting of word-compounding, respectively known as the Sandhi and Vichchhed, in Sanskrit language. Sandhi is an important idea essential to morphological analysis of Sanskrit texts. Sandhi leads to word transformations at word boundaries. The rules of Sandhi formation are well defined but complex, sometimes optional and in some cases, require knowledge about the nature of the words being compounded. Sandhi split or Vichchhed is an even more difficult task given its non uniqueness and context dependence. In this work, we propose the route of formulating the problem as a sequence to sequence prediction task, using modern deep learning techniques. Being the first fully data driven technique, we demonstrate that our model has an accuracy better than the existing methods on multiple standard datasets, despite not using any additional lexical or morphological resources. The code is being made available at https://github.com/IITD-DataScience/Sandhi_Prakarana
title Neural Compound-Word (Sandhi) Generation and Splitting in Sanskrit Language
topic Computation and Language
url https://arxiv.org/abs/2010.12940