1-800-SHARED-TASKS @ NLU of Devanagari Script Languages: Detection of Language, Hate Speech, and Targets using LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Purbey, Jebish, Pullakhandam, Siddartha, Mehreen, Kanwal, Arham, Muhammad, Sharma, Drishti, Srivastava, Ashay, Kadiyala, Ram Mohan Rao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913575684014080
author Purbey, Jebish
Pullakhandam, Siddartha
Mehreen, Kanwal
Arham, Muhammad
Sharma, Drishti
Srivastava, Ashay
Kadiyala, Ram Mohan Rao
author_facet Purbey, Jebish
Pullakhandam, Siddartha
Mehreen, Kanwal
Arham, Muhammad
Sharma, Drishti
Srivastava, Ashay
Kadiyala, Ram Mohan Rao
contents This paper presents a detailed system description of our entry for the CHiPSAL 2025 shared task, focusing on language detection, hate speech identification, and target detection in Devanagari script languages. We experimented with a combination of large language models and their ensembles, including MuRIL, IndicBERT, and Gemma-2, and leveraged unique techniques like focal loss to address challenges in the natural understanding of Devanagari languages, such as multilingual processing and class imbalance. Our approach achieved competitive results across all tasks: F1 of 0.9980, 0.7652, and 0.6804 for Sub-tasks A, B, and C respectively. This work provides insights into the effectiveness of transformer models in tasks with domain-specific and linguistic challenges, as well as areas for potential improvement in future iterations.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06850
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 1-800-SHARED-TASKS @ NLU of Devanagari Script Languages: Detection of Language, Hate Speech, and Targets using LLMs
Purbey, Jebish
Pullakhandam, Siddartha
Mehreen, Kanwal
Arham, Muhammad
Sharma, Drishti
Srivastava, Ashay
Kadiyala, Ram Mohan Rao
Computation and Language
Artificial Intelligence
Machine Learning
This paper presents a detailed system description of our entry for the CHiPSAL 2025 shared task, focusing on language detection, hate speech identification, and target detection in Devanagari script languages. We experimented with a combination of large language models and their ensembles, including MuRIL, IndicBERT, and Gemma-2, and leveraged unique techniques like focal loss to address challenges in the natural understanding of Devanagari languages, such as multilingual processing and class imbalance. Our approach achieved competitive results across all tasks: F1 of 0.9980, 0.7652, and 0.6804 for Sub-tasks A, B, and C respectively. This work provides insights into the effectiveness of transformer models in tasks with domain-specific and linguistic challenges, as well as areas for potential improvement in future iterations.
title 1-800-SHARED-TASKS @ NLU of Devanagari Script Languages: Detection of Language, Hate Speech, and Targets using LLMs
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.06850