Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
Fuente:
arXiv
Saved in:
| Main Authors: | Shimgekar, Soorya Ram, Goyal, Agam, Parulekar, Amruta, Chen, Joshua, Wang, Yian, Kumar, Navin, Sundaram, Hari, Chandrasekharan, Eshwar, Saha, Koustuv |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Social Simulacra in the Wild: AI Agent Communities on Moltbook
by: Goyal, Agam, et al.
Published: (2026)
by: Goyal, Agam, et al.
Published: (2026)
From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities
by: Goyal, Agam, et al.
Published: (2026)
by: Goyal, Agam, et al.
Published: (2026)
Detecting Early and Implicit Suicidal Ideation via Longitudinal and Information Environment Signals on Social Media
by: Shimgekar, Soorya Ram, et al.
Published: (2025)
by: Shimgekar, Soorya Ram, et al.
Published: (2025)
Algorithmic Cultivation: How Social Media Feeds Shape User Language
by: Pal, Olivia, et al.
Published: (2026)
by: Pal, Olivia, et al.
Published: (2026)
The Hidden Toll of Social Media News: Causal Effects on Psychosocial Wellbeing
by: Pal, Olivia, et al.
Published: (2026)
by: Pal, Olivia, et al.
Published: (2026)
Agentic AI framework for End-to-End Medical Data Inference
by: Shimgekar, Soorya Ram, et al.
Published: (2025)
by: Shimgekar, Soorya Ram, et al.
Published: (2025)
Answer Bubbles: Information Exposure in AI-Mediated Search
by: Huang, Michelle, et al.
Published: (2026)
by: Huang, Michelle, et al.
Published: (2026)
LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback
by: Kim, Jiwon, et al.
Published: (2026)
by: Kim, Jiwon, et al.
Published: (2026)
AI Psychosis: Does Conversational AI Amplify Delusion-Related Language?
by: Shimgekar, Soorya Ram, et al.
Published: (2026)
by: Shimgekar, Soorya Ram, et al.
Published: (2026)
SLM-Mod: Small Language Models Surpass LLMs at Content Moderation
by: Zhan, Xianyang, et al.
Published: (2024)
by: Zhan, Xianyang, et al.
Published: (2024)
MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online Governance
by: Goyal, Agam, et al.
Published: (2025)
by: Goyal, Agam, et al.
Published: (2025)
VASTU: Value-Aligned Social Toolkit for Online Content Curation
by: Goyal, Agam, et al.
Published: (2026)
by: Goyal, Agam, et al.
Published: (2026)
ArgCMV: An Argument Summarization Benchmark for the LLM-era
by: Gurjar, Omkar, et al.
Published: (2025)
by: Gurjar, Omkar, et al.
Published: (2025)
The Language of Approval: Identifying the Drivers of Positive Feedback Online
by: Goyal, Agam, et al.
Published: (2025)
by: Goyal, Agam, et al.
Published: (2025)
Does Positive Reinforcement Work?: A Quasi-Experimental Study of the Effects of Positive Feedback on Reddit
by: Lambert, Charlotte, et al.
Published: (2024)
by: Lambert, Charlotte, et al.
Published: (2024)
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification
by: Wang, Yian, et al.
Published: (2026)
by: Wang, Yian, et al.
Published: (2026)
Interpersonal Theory of Suicide as a Lens to Examine Suicidal Ideation in Online Spaces
by: Shimgekar, Soorya Ram, et al.
Published: (2025)
by: Shimgekar, Soorya Ram, et al.
Published: (2025)
Mind Your Ps and Qs: Supporting Positive Reinforcement in Moderation Through a Positive Queue
by: Lambert, Charlotte, et al.
Published: (2025)
by: Lambert, Charlotte, et al.
Published: (2025)
The Ranking Effect: How Algorithmic Rank Influences Attention on Social Media
by: Chan, Jackie, et al.
Published: (2025)
by: Chan, Jackie, et al.
Published: (2025)
Uncovering the Internet's Hidden Values: An Empirical Study of Desirable Behavior Using Highly-Upvoted Content on Reddit
by: Goyal, Agam, et al.
Published: (2024)
by: Goyal, Agam, et al.
Published: (2024)
Examining Algorithmic Curation on Social Media: An Empirical Audit of Reddit's r/popular Feed
by: Chan, Jackie, et al.
Published: (2025)
by: Chan, Jackie, et al.
Published: (2025)
Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders
by: Goyal, Agam, et al.
Published: (2025)
by: Goyal, Agam, et al.
Published: (2025)
PAIR-SAFE: A Paired-Agent Approach for Runtime Auditing and Refining AI-Mediated Mental Health Support
by: Kim, Jiwon, et al.
Published: (2026)
by: Kim, Jiwon, et al.
Published: (2026)
Belief Is All You Need: Modeling Narrative Archetypes in Conspiratorial Discourse
by: Shimgekar, Soorya Ram, et al.
Published: (2025)
by: Shimgekar, Soorya Ram, et al.
Published: (2025)
Do We Know What They Know We Know? Calibrating Student Trust in AI and Human Responses Through Mutual Theory of Mind
by: Pal, Olivia, et al.
Published: (2026)
by: Pal, Olivia, et al.
Published: (2026)
The Chilling: Identifying Strategic Antisocial Behavior Online and Examining the Impact on Journalists
by: Wang, Yian, et al.
Published: (2025)
by: Wang, Yian, et al.
Published: (2025)
LASER: An LLM-based ASR Scoring and Evaluation Rubric
by: Parulekar, Amruta, et al.
Published: (2025)
by: Parulekar, Amruta, et al.
Published: (2025)
Masking or Mitigating? Deconstructing the Impact of Query Rewriting on Retriever Biases in RAG
by: Goyal, Agam, et al.
Published: (2026)
by: Goyal, Agam, et al.
Published: (2026)
Towards a Better Modqueue: Designing for Diversity Across Moderator Objectives and Workflows
by: Bajpai, Tanvi, et al.
Published: (2024)
by: Bajpai, Tanvi, et al.
Published: (2024)
"Think about it like you're a firefighter": Understanding How Reddit Moderators Use the Modqueue
by: Bajpai, Tanvi, et al.
Published: (2025)
by: Bajpai, Tanvi, et al.
Published: (2025)
Designing Usable Controls for Customizable Social Media Feeds
by: Choi, Frederick, et al.
Published: (2025)
by: Choi, Frederick, et al.
Published: (2025)
Venire: A Machine Learning-Guided Panel Review System for Community Content Moderation
by: Koshy, Vinay, et al.
Published: (2024)
by: Koshy, Vinay, et al.
Published: (2024)
From Wearables to Warnings: Predicting Pain Spikes in Patients with Opioid Use Disorder
by: Goyal, Abhay, et al.
Published: (2025)
by: Goyal, Abhay, et al.
Published: (2025)
Needling Through the Threads: A Visualization Tool for Navigating Threaded Online Discussions
by: Liu, Yijun, et al.
Published: (2025)
by: Liu, Yijun, et al.
Published: (2025)
State Contamination in Memory-Augmented LLM Agents
by: Wang, Yian, et al.
Published: (2026)
by: Wang, Yian, et al.
Published: (2026)
The Role of Generative AI in Software Student CollaborAItion
by: Kiesler, Natalie, et al.
Published: (2025)
by: Kiesler, Natalie, et al.
Published: (2025)
One VLM, Two Roles: Stage-Wise Routing and Specialty-Level Deployment for Clinical Workflows
by: Vassef, Shayan, et al.
Published: (2025)
by: Vassef, Shayan, et al.
Published: (2025)
AMPS: ASR with Multimodal Paraphrase Supervision
by: Gupta, Abhishek, et al.
Published: (2024)
by: Gupta, Abhishek, et al.
Published: (2024)
Parameter-efficient Adaptation of Multilingual Multimodal Models for Low-resource ASR
by: Gupta, Abhishek, et al.
Published: (2024)
by: Gupta, Abhishek, et al.
Published: (2024)
Causal Stories from Sensor Traces: Auditing Epistemic Overreach in LLM-Generated Personal Sensing Explanations
by: Zhu, Shanshan, et al.
Published: (2026)
by: Zhu, Shanshan, et al.
Published: (2026)
Similar Items
-
Social Simulacra in the Wild: AI Agent Communities on Moltbook
by: Goyal, Agam, et al.
Published: (2026) -
From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities
by: Goyal, Agam, et al.
Published: (2026) -
Detecting Early and Implicit Suicidal Ideation via Longitudinal and Information Environment Signals on Social Media
by: Shimgekar, Soorya Ram, et al.
Published: (2025) -
Algorithmic Cultivation: How Social Media Feeds Shape User Language
by: Pal, Olivia, et al.
Published: (2026) -
The Hidden Toll of Social Media News: Causal Effects on Psychosocial Wellbeing
by: Pal, Olivia, et al.
Published: (2026)