Analyzing LLM Reasoning to Uncover Mental Health Stigma

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sankar, Sreehari, Nafar, Aliakbar, Barman, Mona, Heitz, Hannah K., Kumar, Ashwin, Tohidi, Pouria, Li, Dailun, Hussain, Danish, DuBois, Russell, Hasheminia, Hamed, Majzoubi, Farshad
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915962111918080
author Sankar, Sreehari
Nafar, Aliakbar
Barman, Mona
Heitz, Hannah K.
Kumar, Ashwin
Tohidi, Pouria
Li, Dailun
Hussain, Danish
DuBois, Russell
Hasheminia, Hamed
Majzoubi, Farshad
author_facet Sankar, Sreehari
Nafar, Aliakbar
Barman, Mona
Heitz, Hannah K.
Kumar, Ashwin
Tohidi, Pouria
Li, Dailun
Hussain, Danish
DuBois, Russell
Hasheminia, Hamed
Majzoubi, Farshad
contents While large language models (LLMs) are increasingly being explored for mental health applications, recent studies reveal that they can exhibit stigma toward individuals with psychological conditions. Existing evaluations of this stigma primarily rely on multiple-choice questions (MCQs), which fail to capture the biases embedded within the models' underlying logic. In this paper, we analyze the intermediate reasoning steps of LLMs to uncover hidden stigmatizing language and the internal rationales driving it. We leverage clinical expertise to categorize common patterns of stigmatizing language directed at individuals with psychological conditions and use this framework to identify and tag problematic statements in LLM reasoning. Furthermore, we rate the severity of these statements, distinguishing between overt prejudice and more subtle, less immediately harmful biases. To broaden the reasoning domain and capture a wider array of patterns, we also extend an existing mental health stigma benchmark by incorporating additional psychological conditions. Our findings demonstrate that evaluating model reasoning not only exposes substantially more stigma than traditional MCQ-based methods but it helps to identify the flaws in the LLMs' logic and their understanding of mental health conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2604_25053
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Analyzing LLM Reasoning to Uncover Mental Health Stigma
Sankar, Sreehari
Nafar, Aliakbar
Barman, Mona
Heitz, Hannah K.
Kumar, Ashwin
Tohidi, Pouria
Li, Dailun
Hussain, Danish
DuBois, Russell
Hasheminia, Hamed
Majzoubi, Farshad
Computation and Language
Artificial Intelligence
I.2.7
While large language models (LLMs) are increasingly being explored for mental health applications, recent studies reveal that they can exhibit stigma toward individuals with psychological conditions. Existing evaluations of this stigma primarily rely on multiple-choice questions (MCQs), which fail to capture the biases embedded within the models' underlying logic. In this paper, we analyze the intermediate reasoning steps of LLMs to uncover hidden stigmatizing language and the internal rationales driving it. We leverage clinical expertise to categorize common patterns of stigmatizing language directed at individuals with psychological conditions and use this framework to identify and tag problematic statements in LLM reasoning. Furthermore, we rate the severity of these statements, distinguishing between overt prejudice and more subtle, less immediately harmful biases. To broaden the reasoning domain and capture a wider array of patterns, we also extend an existing mental health stigma benchmark by incorporating additional psychological conditions. Our findings demonstrate that evaluating model reasoning not only exposes substantially more stigma than traditional MCQ-based methods but it helps to identify the flaws in the LLMs' logic and their understanding of mental health conditions.
title Analyzing LLM Reasoning to Uncover Mental Health Stigma
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2604.25053