Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shimgekar, Soorya Ram, Goyal, Agam, Parulekar, Amruta, Chen, Joshua, Wang, Yian, Kumar, Navin, Sundaram, Hari, Chandrasekharan, Eshwar, Saha, Koustuv
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916064538918912
author Shimgekar, Soorya Ram
Goyal, Agam
Parulekar, Amruta
Chen, Joshua
Wang, Yian
Kumar, Navin
Sundaram, Hari
Chandrasekharan, Eshwar
Saha, Koustuv
author_facet Shimgekar, Soorya Ram
Goyal, Agam
Parulekar, Amruta
Chen, Joshua
Wang, Yian
Kumar, Navin
Sundaram, Hari
Chandrasekharan, Eshwar
Saha, Koustuv
contents Large language models (LLMs) are increasingly deployed in conversational settings where user tone ranges from polite to adversarial or toxic, yet less is known about whether toxic language in otherwise semantically equivalent prompts can degrade factual reliability. We study how lexical and tone-based prompt perturbations affect the factual reliability of LLMs. Using controlled prompt variations across polite, random, and three toxicity levels, we evaluate five LLMs on ARC-Easy, GSM8K, and MMLU. We find that toxic lexical perturbations consistently reduce factual accuracy and increase uncertainty, while polite phrasing yields limited and inconsistent changes. To examine whether these answer inconsistencies correspond to internal changes, we conduct attribution-graph analyses of model activations and influences. We find that increasing toxicity selectively amplifies perturbation-sensitive variant nodes while relatively stable core reasoning nodes remain more invariant. These findings position prompt tone as a critical dimension of LLM reliability and provide behavioral and mechanistic evidence that surface-level lexical variation can alter factual outputs and internal computation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30913
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
Shimgekar, Soorya Ram
Goyal, Agam
Parulekar, Amruta
Chen, Joshua
Wang, Yian
Kumar, Navin
Sundaram, Hari
Chandrasekharan, Eshwar
Saha, Koustuv
Computation and Language
Artificial Intelligence
Computers and Society
Human-Computer Interaction
Large language models (LLMs) are increasingly deployed in conversational settings where user tone ranges from polite to adversarial or toxic, yet less is known about whether toxic language in otherwise semantically equivalent prompts can degrade factual reliability. We study how lexical and tone-based prompt perturbations affect the factual reliability of LLMs. Using controlled prompt variations across polite, random, and three toxicity levels, we evaluate five LLMs on ARC-Easy, GSM8K, and MMLU. We find that toxic lexical perturbations consistently reduce factual accuracy and increase uncertainty, while polite phrasing yields limited and inconsistent changes. To examine whether these answer inconsistencies correspond to internal changes, we conduct attribution-graph analyses of model activations and influences. We find that increasing toxicity selectively amplifies perturbation-sensitive variant nodes while relatively stable core reasoning nodes remain more invariant. These findings position prompt tone as a critical dimension of LLM reliability and provide behavioral and mechanistic evidence that surface-level lexical variation can alter factual outputs and internal computation.
title Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
topic Computation and Language
Artificial Intelligence
Computers and Society
Human-Computer Interaction
url https://arxiv.org/abs/2605.30913