Toxicity Ahead: Forecasting Conversational Derailment on GitHub

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Imran, Mia Mohammad, Zita, Robert, Rahman, Rahat Rizvi, Chatterjee, Preetha, Damevski, Kostadin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918251855872000
author Imran, Mia Mohammad
Zita, Robert
Rahman, Rahat Rizvi
Chatterjee, Preetha
Damevski, Kostadin
author_facet Imran, Mia Mohammad
Zita, Robert
Rahman, Rahat Rizvi
Chatterjee, Preetha
Damevski, Kostadin
contents Toxic interactions in Open Source Software (OSS) communities reduce contributor engagement and threaten project sustainability. Preventing such toxicity before it emerges requires a clear understanding of how harmful conversations unfold. However, most proactive moderation strategies are manual, requiring significant time and effort from community maintainers. To support more scalable approaches, we curate a dataset of 159 derailed toxic threads and 207 non-toxic threads from GitHub discussions. Our analysis reveals that toxicity can be forecast by tension triggers, sentiment shifts, and specific conversational patterns. We present a novel Large Language Model (LLM)-based framework for predicting conversational derailment on GitHub using a two-step prompting pipeline. First, we generate \textit{Summaries of Conversation Dynamics} (SCDs) via Least-to-Most (LtM) prompting; then we use these summaries to estimate the \textit{likelihood of derailment}. Evaluated on Qwen and Llama models, our LtM strategy achieves F1-scores of 0.901 and 0.852, respectively, at a decision threshold of 0.3, outperforming established NLP baselines on conversation derailment. External validation on a dataset of 308 GitHub issue threads (65 toxic, 243 non-toxic) yields an F1-score up to 0.797. Our findings demonstrate the effectiveness of structured LLM prompting for early detection of conversational derailment in OSS, enabling proactive and explainable moderation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_15031
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Toxicity Ahead: Forecasting Conversational Derailment on GitHub
Imran, Mia Mohammad
Zita, Robert
Rahman, Rahat Rizvi
Chatterjee, Preetha
Damevski, Kostadin
Software Engineering
Computers and Society
Human-Computer Interaction
Toxic interactions in Open Source Software (OSS) communities reduce contributor engagement and threaten project sustainability. Preventing such toxicity before it emerges requires a clear understanding of how harmful conversations unfold. However, most proactive moderation strategies are manual, requiring significant time and effort from community maintainers. To support more scalable approaches, we curate a dataset of 159 derailed toxic threads and 207 non-toxic threads from GitHub discussions. Our analysis reveals that toxicity can be forecast by tension triggers, sentiment shifts, and specific conversational patterns. We present a novel Large Language Model (LLM)-based framework for predicting conversational derailment on GitHub using a two-step prompting pipeline. First, we generate \textit{Summaries of Conversation Dynamics} (SCDs) via Least-to-Most (LtM) prompting; then we use these summaries to estimate the \textit{likelihood of derailment}. Evaluated on Qwen and Llama models, our LtM strategy achieves F1-scores of 0.901 and 0.852, respectively, at a decision threshold of 0.3, outperforming established NLP baselines on conversation derailment. External validation on a dataset of 308 GitHub issue threads (65 toxic, 243 non-toxic) yields an F1-score up to 0.797. Our findings demonstrate the effectiveness of structured LLM prompting for early detection of conversational derailment in OSS, enabling proactive and explainable moderation.
title Toxicity Ahead: Forecasting Conversational Derailment on GitHub
topic Software Engineering
Computers and Society
Human-Computer Interaction
url https://arxiv.org/abs/2512.15031