User-Aware Multilingual Abusive Content Detection in Social Media

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rehman, Mohammad Zia Ur, Mehta, Somya, Singh, Kuldeep, Kaushik, Kunal, Kumar, Nagendra
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909368863162368
author Rehman, Mohammad Zia Ur
Mehta, Somya
Singh, Kuldeep
Kaushik, Kunal
Kumar, Nagendra
author_facet Rehman, Mohammad Zia Ur
Mehta, Somya
Singh, Kuldeep
Kaushik, Kunal
Kumar, Nagendra
contents Despite growing efforts to halt distasteful content on social media, multilingualism has added a new dimension to this problem. The scarcity of resources makes the challenge even greater when it comes to low-resource languages. This work focuses on providing a novel method for abusive content detection in multiple low-resource Indic languages. Our observation indicates that a post's tendency to attract abusive comments, as well as features such as user history and social context, significantly aid in the detection of abusive content. The proposed method first learns social and text context features in two separate modules. The integrated representation from these modules is learned and used for the final prediction. To evaluate the performance of our method against different classical and state-of-the-art methods, we have performed extensive experiments on SCIDN and MACI datasets consisting of 1.5M and 665K multilingual comments, respectively. Our proposed method outperforms state-of-the-art baseline methods with an average increase of 4.08% and 9.52% in F1-scores on SCIDN and MACI datasets, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21321
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle User-Aware Multilingual Abusive Content Detection in Social Media
Rehman, Mohammad Zia Ur
Mehta, Somya
Singh, Kuldeep
Kaushik, Kunal
Kumar, Nagendra
Social and Information Networks
Artificial Intelligence
Computation and Language
Despite growing efforts to halt distasteful content on social media, multilingualism has added a new dimension to this problem. The scarcity of resources makes the challenge even greater when it comes to low-resource languages. This work focuses on providing a novel method for abusive content detection in multiple low-resource Indic languages. Our observation indicates that a post's tendency to attract abusive comments, as well as features such as user history and social context, significantly aid in the detection of abusive content. The proposed method first learns social and text context features in two separate modules. The integrated representation from these modules is learned and used for the final prediction. To evaluate the performance of our method against different classical and state-of-the-art methods, we have performed extensive experiments on SCIDN and MACI datasets consisting of 1.5M and 665K multilingual comments, respectively. Our proposed method outperforms state-of-the-art baseline methods with an average increase of 4.08% and 9.52% in F1-scores on SCIDN and MACI datasets, respectively.
title User-Aware Multilingual Abusive Content Detection in Social Media
topic Social and Information Networks
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.21321