UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912464253222912 |
|---|---|
| author | Beniwal, Himanshu Venkat, Reddybathuni Kumar, Rohit Srivibhav, Birudugadda Jain, Daksh Doddi, Pavan Dhande, Eshwar Ananth, Adithya Kuldeep Singh, Mayank |
| author_facet | Beniwal, Himanshu Venkat, Reddybathuni Kumar, Rohit Srivibhav, Birudugadda Jain, Daksh Doddi, Pavan Dhande, Eshwar Ananth, Adithya Kuldeep Singh, Mayank |
| contents | This work introduces UnityAI-Guard, a framework for binary toxicity classification targeting low-resource Indian languages. While existing systems predominantly cater to high-resource languages, UnityAI-Guard addresses this critical gap by developing state-of-the-art models for identifying toxic content across diverse Brahmic/Indic scripts. Our approach achieves an impressive average F1-score of 84.23% across seven languages, leveraging a dataset of 567k training instances and 30k manually verified test instances. By advancing multilingual content moderation for linguistically diverse regions, UnityAI-Guard also provides public API access to foster broader adoption and application. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_23088 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages Beniwal, Himanshu Venkat, Reddybathuni Kumar, Rohit Srivibhav, Birudugadda Jain, Daksh Doddi, Pavan Dhande, Eshwar Ananth, Adithya Kuldeep Singh, Mayank Computation and Language Artificial Intelligence This work introduces UnityAI-Guard, a framework for binary toxicity classification targeting low-resource Indian languages. While existing systems predominantly cater to high-resource languages, UnityAI-Guard addresses this critical gap by developing state-of-the-art models for identifying toxic content across diverse Brahmic/Indic scripts. Our approach achieves an impressive average F1-score of 84.23% across seven languages, leveraging a dataset of 567k training instances and 30k manually verified test instances. By advancing multilingual content moderation for linguistically diverse regions, UnityAI-Guard also provides public API access to foster broader adoption and application. |
| title | UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2503.23088 |