Deep Language Detection for Indian Code: A Context-Aware Deep Learning for Multilingual Content Classification

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autori principali: Yadav, Shraddha, Srivastava, Esha, Srivastava, Amit
Natura: Recurso digital
Lingua:inglese
Pubblicazione: Zenodo 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866901044560134144
author Yadav, Shraddha
Srivastava, Esha
Srivastava, Amit
author_facet Yadav, Shraddha
Srivastava, Esha
Srivastava, Amit
contents <p>The multilingual character of India has been seen in the online discussions whereby Hindi, gujarati, and English are highly mixed in a single</p> <p>sentence. Such a phenomenon is referred to as code-mixing and it is quite challenging to standard Natural Language Processing (NLP) systems,</p> <p>especially language identification. Conventional models, which most often were trained on monolingual, standardized data, cannot cope with</p> <p>informal, transliterated, or script-vulnerable text that is frequent on the social media. The proposed paper suggests a deep learning-based</p> <p>language detection model that is tailored to Indian code-mixed text. The model combines context-sensitive word representations of multilingual</p> <p>transformers (mBERT) and bidirectional LSTM sequence decoders with attention mechanisms to identify and tag of languages at the word and</p> <p>sentence levels. A linguistically annotated, manually curated set of Hindi-English-Gujarati Reddit and Twitter posts was constructed and</p> <p>annotated. Experiments show that the suggested model significantly performs better than the baseline mechanisms, especially when it comes to</p> <p>highly mixed and informal content. Moreover, a script-sensitive pre-processing pipeline improves the detection of Roman, Devanagari and</p> <p>Gujarati scripts. The results are promising to the creation of inclusive language technologies in India, such as a chatbot or content moderation</p> <p>system and multilingual information retrieval. This study addresses the problems of low-resource and code-mixed language processing and thus</p> <p>bridges the gap in the research on Indian NLP and preconditions further progress in the field of working with complex, multilingual, and</p> <p>informal online text.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18107005
institution Zenodo
language eng
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle Deep Language Detection for Indian Code: A Context-Aware Deep Learning for Multilingual Content Classification
Yadav, Shraddha
Srivastava, Esha
Srivastava, Amit
Multilingual Language Detection
Hindi–Gujarati Code-Mixed Text
Natural Language Processing (NLP)
Deep Learning,
Transformer Models
IndicBERT
BERT
BiLSTM
Self-Attention
Conditional Random Fields (CRF)
Transliteration Challenges
<p>The multilingual character of India has been seen in the online discussions whereby Hindi, gujarati, and English are highly mixed in a single</p> <p>sentence. Such a phenomenon is referred to as code-mixing and it is quite challenging to standard Natural Language Processing (NLP) systems,</p> <p>especially language identification. Conventional models, which most often were trained on monolingual, standardized data, cannot cope with</p> <p>informal, transliterated, or script-vulnerable text that is frequent on the social media. The proposed paper suggests a deep learning-based</p> <p>language detection model that is tailored to Indian code-mixed text. The model combines context-sensitive word representations of multilingual</p> <p>transformers (mBERT) and bidirectional LSTM sequence decoders with attention mechanisms to identify and tag of languages at the word and</p> <p>sentence levels. A linguistically annotated, manually curated set of Hindi-English-Gujarati Reddit and Twitter posts was constructed and</p> <p>annotated. Experiments show that the suggested model significantly performs better than the baseline mechanisms, especially when it comes to</p> <p>highly mixed and informal content. Moreover, a script-sensitive pre-processing pipeline improves the detection of Roman, Devanagari and</p> <p>Gujarati scripts. The results are promising to the creation of inclusive language technologies in India, such as a chatbot or content moderation</p> <p>system and multilingual information retrieval. This study addresses the problems of low-resource and code-mixed language processing and thus</p> <p>bridges the gap in the research on Indian NLP and preconditions further progress in the field of working with complex, multilingual, and</p> <p>informal online text.</p>
title Deep Language Detection for Indian Code: A Context-Aware Deep Learning for Multilingual Content Classification
topic Multilingual Language Detection
Hindi–Gujarati Code-Mixed Text
Natural Language Processing (NLP)
Deep Learning,
Transformer Models
IndicBERT
BERT
BiLSTM
Self-Attention
Conditional Random Fields (CRF)
Transliteration Challenges
url https://doi.org/10.5281/zenodo.18107005