UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Beniwal, Himanshu, Venkat, Reddybathuni, Kumar, Rohit, Srivibhav, Birudugadda, Jain, Daksh, Doddi, Pavan, Dhande, Eshwar, Ananth, Adithya, Kuldeep, Singh, Mayank
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912464253222912
author Beniwal, Himanshu
Venkat, Reddybathuni
Kumar, Rohit
Srivibhav, Birudugadda
Jain, Daksh
Doddi, Pavan
Dhande, Eshwar
Ananth, Adithya
Kuldeep
Singh, Mayank
author_facet Beniwal, Himanshu
Venkat, Reddybathuni
Kumar, Rohit
Srivibhav, Birudugadda
Jain, Daksh
Doddi, Pavan
Dhande, Eshwar
Ananth, Adithya
Kuldeep
Singh, Mayank
contents This work introduces UnityAI-Guard, a framework for binary toxicity classification targeting low-resource Indian languages. While existing systems predominantly cater to high-resource languages, UnityAI-Guard addresses this critical gap by developing state-of-the-art models for identifying toxic content across diverse Brahmic/Indic scripts. Our approach achieves an impressive average F1-score of 84.23% across seven languages, leveraging a dataset of 567k training instances and 30k manually verified test instances. By advancing multilingual content moderation for linguistically diverse regions, UnityAI-Guard also provides public API access to foster broader adoption and application.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23088
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
Beniwal, Himanshu
Venkat, Reddybathuni
Kumar, Rohit
Srivibhav, Birudugadda
Jain, Daksh
Doddi, Pavan
Dhande, Eshwar
Ananth, Adithya
Kuldeep
Singh, Mayank
Computation and Language
Artificial Intelligence
This work introduces UnityAI-Guard, a framework for binary toxicity classification targeting low-resource Indian languages. While existing systems predominantly cater to high-resource languages, UnityAI-Guard addresses this critical gap by developing state-of-the-art models for identifying toxic content across diverse Brahmic/Indic scripts. Our approach achieves an impressive average F1-score of 84.23% across seven languages, leveraging a dataset of 567k training instances and 30k manually verified test instances. By advancing multilingual content moderation for linguistically diverse regions, UnityAI-Guard also provides public API access to foster broader adoption and application.
title UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.23088