Deep Language Detection for Indian Code: A Context-Aware Deep Learning for Multilingual Content Classification
Fuente:
Zenodo
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Recurso digital |
| Lingua: | inglese |
| Pubblicazione: |
Zenodo
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866901044560134144 |
|---|---|
| author | Yadav, Shraddha Srivastava, Esha Srivastava, Amit |
| author_facet | Yadav, Shraddha Srivastava, Esha Srivastava, Amit |
| contents | <p>The multilingual character of India has been seen in the online discussions whereby Hindi, gujarati, and English are highly mixed in a single</p> <p>sentence. Such a phenomenon is referred to as code-mixing and it is quite challenging to standard Natural Language Processing (NLP) systems,</p> <p>especially language identification. Conventional models, which most often were trained on monolingual, standardized data, cannot cope with</p> <p>informal, transliterated, or script-vulnerable text that is frequent on the social media. The proposed paper suggests a deep learning-based</p> <p>language detection model that is tailored to Indian code-mixed text. The model combines context-sensitive word representations of multilingual</p> <p>transformers (mBERT) and bidirectional LSTM sequence decoders with attention mechanisms to identify and tag of languages at the word and</p> <p>sentence levels. A linguistically annotated, manually curated set of Hindi-English-Gujarati Reddit and Twitter posts was constructed and</p> <p>annotated. Experiments show that the suggested model significantly performs better than the baseline mechanisms, especially when it comes to</p> <p>highly mixed and informal content. Moreover, a script-sensitive pre-processing pipeline improves the detection of Roman, Devanagari and</p> <p>Gujarati scripts. The results are promising to the creation of inclusive language technologies in India, such as a chatbot or content moderation</p> <p>system and multilingual information retrieval. This study addresses the problems of low-resource and code-mixed language processing and thus</p> <p>bridges the gap in the research on Indian NLP and preconditions further progress in the field of working with complex, multilingual, and</p> <p>informal online text.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18107005 |
| institution | Zenodo |
| language | eng |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Deep Language Detection for Indian Code: A Context-Aware Deep Learning for Multilingual Content Classification Yadav, Shraddha Srivastava, Esha Srivastava, Amit Multilingual Language Detection Hindi–Gujarati Code-Mixed Text Natural Language Processing (NLP) Deep Learning, Transformer Models IndicBERT BERT BiLSTM Self-Attention Conditional Random Fields (CRF) Transliteration Challenges <p>The multilingual character of India has been seen in the online discussions whereby Hindi, gujarati, and English are highly mixed in a single</p> <p>sentence. Such a phenomenon is referred to as code-mixing and it is quite challenging to standard Natural Language Processing (NLP) systems,</p> <p>especially language identification. Conventional models, which most often were trained on monolingual, standardized data, cannot cope with</p> <p>informal, transliterated, or script-vulnerable text that is frequent on the social media. The proposed paper suggests a deep learning-based</p> <p>language detection model that is tailored to Indian code-mixed text. The model combines context-sensitive word representations of multilingual</p> <p>transformers (mBERT) and bidirectional LSTM sequence decoders with attention mechanisms to identify and tag of languages at the word and</p> <p>sentence levels. A linguistically annotated, manually curated set of Hindi-English-Gujarati Reddit and Twitter posts was constructed and</p> <p>annotated. Experiments show that the suggested model significantly performs better than the baseline mechanisms, especially when it comes to</p> <p>highly mixed and informal content. Moreover, a script-sensitive pre-processing pipeline improves the detection of Roman, Devanagari and</p> <p>Gujarati scripts. The results are promising to the creation of inclusive language technologies in India, such as a chatbot or content moderation</p> <p>system and multilingual information retrieval. This study addresses the problems of low-resource and code-mixed language processing and thus</p> <p>bridges the gap in the research on Indian NLP and preconditions further progress in the field of working with complex, multilingual, and</p> <p>informal online text.</p> |
| title | Deep Language Detection for Indian Code: A Context-Aware Deep Learning for Multilingual Content Classification |
| topic | Multilingual Language Detection Hindi–Gujarati Code-Mixed Text Natural Language Processing (NLP) Deep Learning, Transformer Models IndicBERT BERT BiLSTM Self-Attention Conditional Random Fields (CRF) Transliteration Challenges |
| url | https://doi.org/10.5281/zenodo.18107005 |