What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ferrara, Alfio, Picascia, Sergio, Pinnavaia, Laura, Ranitovic, Vojimir, Rocchetti, Elisabetta, Tuveri, Alice
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912511969722368
author Ferrara, Alfio
Picascia, Sergio
Pinnavaia, Laura
Ranitovic, Vojimir
Rocchetti, Elisabetta
Tuveri, Alice
author_facet Ferrara, Alfio
Picascia, Sergio
Pinnavaia, Laura
Ranitovic, Vojimir
Rocchetti, Elisabetta
Tuveri, Alice
contents Proprietary Large Language Models (LLMs) have shown tendencies toward politeness, formality, and implicit content moderation. While previous research has primarily focused on explicitly training models to moderate and detoxify sensitive content, there has been limited exploration of whether LLMs implicitly sanitize language without explicit instructions. This study empirically analyzes the implicit moderation behavior of GPT-4o-mini when paraphrasing sensitive content and evaluates the extent of sensitivity shifts. Our experiments indicate that GPT-4o-mini systematically moderates content toward less sensitive classes, with substantial reductions in derogatory and taboo language. Also, we evaluate the zero-shot capabilities of LLMs in classifying sentence sensitivity, comparing their performances against traditional methods.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23319
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content
Ferrara, Alfio
Picascia, Sergio
Pinnavaia, Laura
Ranitovic, Vojimir
Rocchetti, Elisabetta
Tuveri, Alice
Computation and Language
Proprietary Large Language Models (LLMs) have shown tendencies toward politeness, formality, and implicit content moderation. While previous research has primarily focused on explicitly training models to moderate and detoxify sensitive content, there has been limited exploration of whether LLMs implicitly sanitize language without explicit instructions. This study empirically analyzes the implicit moderation behavior of GPT-4o-mini when paraphrasing sensitive content and evaluates the extent of sensitivity shifts. Our experiments indicate that GPT-4o-mini systematically moderates content toward less sensitive classes, with substantial reductions in derogatory and taboo language. Also, we evaluate the zero-shot capabilities of LLMs in classifying sentence sensitivity, comparing their performances against traditional methods.
title What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content
topic Computation and Language
url https://arxiv.org/abs/2507.23319