Real-Time Toxicity Filtering for Open-Source Code Reviews

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Anindya, Md Awsaf Alam, Biswas, Showvik, Iqbal, Anindya, Sarker, Jaydeb, Bosu, Amiangshu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908951067492352
author Anindya, Md Awsaf Alam
Biswas, Showvik
Iqbal, Anindya
Sarker, Jaydeb
Bosu, Amiangshu
author_facet Anindya, Md Awsaf Alam
Biswas, Showvik
Iqbal, Anindya
Sarker, Jaydeb
Bosu, Amiangshu
contents Toxic interactions in open-source software development harm community collaboration. To combat this, we propose ToxiShield, a realtime browser extension that identifies and detoxifies toxic code reviews. The framework comprises three modules: toxicity identification, reasoned multiclass classification, and code review detoxification. Our fine-tuned BERT-based binary classifier achieved a 97% F1-score on 38,761 code review texts. For multiclass classification, Claude 3.5 Sonnet with prompt engineering achieved a 39% MCC and 42% F1 on 1,200 samples. Finally, our fine-tuned Llama 3.2 detoxification model reached 95.27% style transfer accuracy, 97.03% fluency, 67.07% content preservation, and an 84% J-score. Validation with 10 software developers suggests ToxiShield effectively fosters a more inclusive open-source environment.
format Preprint
id arxiv_https___arxiv_org_abs_2604_08886
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Real-Time Toxicity Filtering for Open-Source Code Reviews
Anindya, Md Awsaf Alam
Biswas, Showvik
Iqbal, Anindya
Sarker, Jaydeb
Bosu, Amiangshu
Software Engineering
Toxic interactions in open-source software development harm community collaboration. To combat this, we propose ToxiShield, a realtime browser extension that identifies and detoxifies toxic code reviews. The framework comprises three modules: toxicity identification, reasoned multiclass classification, and code review detoxification. Our fine-tuned BERT-based binary classifier achieved a 97% F1-score on 38,761 code review texts. For multiclass classification, Claude 3.5 Sonnet with prompt engineering achieved a 39% MCC and 42% F1 on 1,200 samples. Finally, our fine-tuned Llama 3.2 detoxification model reached 95.27% style transfer accuracy, 97.03% fluency, 67.07% content preservation, and an 84% J-score. Validation with 10 software developers suggests ToxiShield effectively fosters a more inclusive open-source environment.
title Real-Time Toxicity Filtering for Open-Source Code Reviews
topic Software Engineering
url https://arxiv.org/abs/2604.08886