Do You Feel Comfortable? Detecting Hidden Conversational Escalation in AI Chatbots

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Jihyung, Afroogh, Saleh, Atkinson, David, Jiao, Junfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909997417365504
author Park, Jihyung
Afroogh, Saleh
Atkinson, David
Jiao, Junfeng
author_facet Park, Jihyung
Afroogh, Saleh
Atkinson, David
Jiao, Junfeng
contents Large Language Models (LLM) are increasingly integrated into everyday interactions, serving not only as information assistants but also as emotional companions. Even in the absence of explicit toxicity, repeated emotional reinforcement or affective drift can gradually escalate distress in a form of \textit{implicit harm} that traditional toxicity filters fail to detect. Existing guardrail mechanisms often rely on external classifiers or clinical rubrics that may lag behind the nuanced, real-time dynamics of a developing conversation. To address this gap, we propose GAUGE (Guarding Affective Utterance Generation Escalation), logit-based framework for the real-time detection of hidden conversational escalation. GAUGE measures how an LLM's output probabilistically shifts the affective state of a dialogue.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06193
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do You Feel Comfortable? Detecting Hidden Conversational Escalation in AI Chatbots
Park, Jihyung
Afroogh, Saleh
Atkinson, David
Jiao, Junfeng
Computation and Language
Artificial Intelligence
Large Language Models (LLM) are increasingly integrated into everyday interactions, serving not only as information assistants but also as emotional companions. Even in the absence of explicit toxicity, repeated emotional reinforcement or affective drift can gradually escalate distress in a form of \textit{implicit harm} that traditional toxicity filters fail to detect. Existing guardrail mechanisms often rely on external classifiers or clinical rubrics that may lag behind the nuanced, real-time dynamics of a developing conversation. To address this gap, we propose GAUGE (Guarding Affective Utterance Generation Escalation), logit-based framework for the real-time detection of hidden conversational escalation. GAUGE measures how an LLM's output probabilistically shifts the affective state of a dialogue.
title Do You Feel Comfortable? Detecting Hidden Conversational Escalation in AI Chatbots
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.06193