Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916064538918912 |
|---|---|
| author | Shimgekar, Soorya Ram Goyal, Agam Parulekar, Amruta Chen, Joshua Wang, Yian Kumar, Navin Sundaram, Hari Chandrasekharan, Eshwar Saha, Koustuv |
| author_facet | Shimgekar, Soorya Ram Goyal, Agam Parulekar, Amruta Chen, Joshua Wang, Yian Kumar, Navin Sundaram, Hari Chandrasekharan, Eshwar Saha, Koustuv |
| contents | Large language models (LLMs) are increasingly deployed in conversational settings where user tone ranges from polite to adversarial or toxic, yet less is known about whether toxic language in otherwise semantically equivalent prompts can degrade factual reliability. We study how lexical and tone-based prompt perturbations affect the factual reliability of LLMs. Using controlled prompt variations across polite, random, and three toxicity levels, we evaluate five LLMs on ARC-Easy, GSM8K, and MMLU. We find that toxic lexical perturbations consistently reduce factual accuracy and increase uncertainty, while polite phrasing yields limited and inconsistent changes. To examine whether these answer inconsistencies correspond to internal changes, we conduct attribution-graph analyses of model activations and influences. We find that increasing toxicity selectively amplifies perturbation-sensitive variant nodes while relatively stable core reasoning nodes remain more invariant. These findings position prompt tone as a critical dimension of LLM reliability and provide behavioral and mechanistic evidence that surface-level lexical variation can alter factual outputs and internal computation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_30913 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits Shimgekar, Soorya Ram Goyal, Agam Parulekar, Amruta Chen, Joshua Wang, Yian Kumar, Navin Sundaram, Hari Chandrasekharan, Eshwar Saha, Koustuv Computation and Language Artificial Intelligence Computers and Society Human-Computer Interaction Large language models (LLMs) are increasingly deployed in conversational settings where user tone ranges from polite to adversarial or toxic, yet less is known about whether toxic language in otherwise semantically equivalent prompts can degrade factual reliability. We study how lexical and tone-based prompt perturbations affect the factual reliability of LLMs. Using controlled prompt variations across polite, random, and three toxicity levels, we evaluate five LLMs on ARC-Easy, GSM8K, and MMLU. We find that toxic lexical perturbations consistently reduce factual accuracy and increase uncertainty, while polite phrasing yields limited and inconsistent changes. To examine whether these answer inconsistencies correspond to internal changes, we conduct attribution-graph analyses of model activations and influences. We find that increasing toxicity selectively amplifies perturbation-sensitive variant nodes while relatively stable core reasoning nodes remain more invariant. These findings position prompt tone as a critical dimension of LLM reliability and provide behavioral and mechanistic evidence that surface-level lexical variation can alter factual outputs and internal computation. |
| title | Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits |
| topic | Computation and Language Artificial Intelligence Computers and Society Human-Computer Interaction |
| url | https://arxiv.org/abs/2605.30913 |