AI Content Moderation in Therapy Conversations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Jiwon, Wang, Claire, Yoon, Taeung, Huang, Sabelle, Saha, Koustuv
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914599242039296
author Kim, Jiwon
Wang, Claire
Yoon, Taeung
Huang, Sabelle
Saha, Koustuv
author_facet Kim, Jiwon
Wang, Claire
Yoon, Taeung
Huang, Sabelle
Saha, Koustuv
contents Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LLMs like ChaptGPT or Llama are often developed with content moderation guardrails that prevent them from discussing sensitive subjects with users for both liability and safety purposes, and this inability to broach these subjects may affect their capacity as therapists. In this study, we perform an algorithm audit on three state-of-the-art moderation systems (OpenAI's moderation endpoint, Meta's Llama Guard, and Google's Shield Gemma) to investigate the extent to which these systems flag the content of real-life therapy sessions as undesirable. Our results raise implications for the limitations that users and organizations may encounter when designing LLMs to play the part of a therapist.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25454
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AI Content Moderation in Therapy Conversations
Kim, Jiwon
Wang, Claire
Yoon, Taeung
Huang, Sabelle
Saha, Koustuv
Human-Computer Interaction
Artificial Intelligence
Computation and Language
Computers and Society
Social and Information Networks
Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LLMs like ChaptGPT or Llama are often developed with content moderation guardrails that prevent them from discussing sensitive subjects with users for both liability and safety purposes, and this inability to broach these subjects may affect their capacity as therapists. In this study, we perform an algorithm audit on three state-of-the-art moderation systems (OpenAI's moderation endpoint, Meta's Llama Guard, and Google's Shield Gemma) to investigate the extent to which these systems flag the content of real-life therapy sessions as undesirable. Our results raise implications for the limitations that users and organizations may encounter when designing LLMs to play the part of a therapist.
title AI Content Moderation in Therapy Conversations
topic Human-Computer Interaction
Artificial Intelligence
Computation and Language
Computers and Society
Social and Information Networks
url https://arxiv.org/abs/2605.25454