Semantically Cohesive Word Grouping in Indian Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Karthika, N J, Patra, Adyasha, Naidu, Nagasai Saketh, Bhattacharya, Arnab, Ramakrishnan, Ganesh, Dangarikar, Chaitali
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912179324715008
author Karthika, N J
Patra, Adyasha
Naidu, Nagasai Saketh
Bhattacharya, Arnab
Ramakrishnan, Ganesh
Dangarikar, Chaitali
author_facet Karthika, N J
Patra, Adyasha
Naidu, Nagasai Saketh
Bhattacharya, Arnab
Ramakrishnan, Ganesh
Dangarikar, Chaitali
contents Indian languages are inflectional and agglutinative and typically follow clause-free word order. The structure of sentences across most major Indian languages are similar when their dependency parse trees are considered. While some differences in the parsing structure occur due to peculiarities of a language or its preferred natural way of conveying meaning, several apparent differences are simply due to the granularity of representation of the smallest semantic unit of processing in a sentence. The semantic unit is typically a word, typographically separated by whitespaces. A single whitespace-separated word in one language may correspond to a group of words in another. Hence, grouping of words based on semantics helps unify the parsing structure of parallel sentences across languages and, in the process, morphology. In this work, we propose word grouping as a major preprocessing step for any computational or linguistic processing of sentences for Indian languages. Among Indian languages, since Hindi is one of the least agglutinative, we expect it to benefit the most from word-grouping. Hence, in this paper, we focus on Hindi to study the effects of grouping. We perform quantitative assessment of our proposal with an intrinsic method that perturbs sentences by shuffling words as well as an extrinsic evaluation that verifies the importance of word grouping for the task of Machine Translation (MT) using decomposed prompting. We also qualitatively analyze certain aspects of the syntactic structure of sentences. Our experiments and analyses show that the proposed grouping technique brings uniformity in the syntactic structures, as well as aids underlying NLP tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03988
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semantically Cohesive Word Grouping in Indian Languages
Karthika, N J
Patra, Adyasha
Naidu, Nagasai Saketh
Bhattacharya, Arnab
Ramakrishnan, Ganesh
Dangarikar, Chaitali
Computation and Language
Indian languages are inflectional and agglutinative and typically follow clause-free word order. The structure of sentences across most major Indian languages are similar when their dependency parse trees are considered. While some differences in the parsing structure occur due to peculiarities of a language or its preferred natural way of conveying meaning, several apparent differences are simply due to the granularity of representation of the smallest semantic unit of processing in a sentence. The semantic unit is typically a word, typographically separated by whitespaces. A single whitespace-separated word in one language may correspond to a group of words in another. Hence, grouping of words based on semantics helps unify the parsing structure of parallel sentences across languages and, in the process, morphology. In this work, we propose word grouping as a major preprocessing step for any computational or linguistic processing of sentences for Indian languages. Among Indian languages, since Hindi is one of the least agglutinative, we expect it to benefit the most from word-grouping. Hence, in this paper, we focus on Hindi to study the effects of grouping. We perform quantitative assessment of our proposal with an intrinsic method that perturbs sentences by shuffling words as well as an extrinsic evaluation that verifies the importance of word grouping for the task of Machine Translation (MT) using decomposed prompting. We also qualitatively analyze certain aspects of the syntactic structure of sentences. Our experiments and analyses show that the proposed grouping technique brings uniformity in the syntactic structures, as well as aids underlying NLP tasks.
title Semantically Cohesive Word Grouping in Indian Languages
topic Computation and Language
url https://arxiv.org/abs/2501.03988