When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Badawi, Abeer, Rahimi, Elahe, Laskar, Md Tahmid Rahman, Grach, Sheri, Bertrand, Lindsay, Danok, Lames, Huang, Jimmy, Rudzicz, Frank, Dolatabadi, Elham
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912663678746624
author Badawi, Abeer
Rahimi, Elahe
Laskar, Md Tahmid Rahman
Grach, Sheri
Bertrand, Lindsay
Danok, Lames
Huang, Jimmy
Rudzicz, Frank
Dolatabadi, Elham
author_facet Badawi, Abeer
Rahimi, Elahe
Laskar, Md Tahmid Rahman
Grach, Sheri
Bertrand, Lindsay
Danok, Lames
Huang, Jimmy
Rudzicz, Frank
Dolatabadi, Elham
contents Evaluating Large Language Models (LLMs) for mental health support is challenging due to the emotionally and cognitively complex nature of therapeutic dialogue. Existing benchmarks are limited in scale, reliability, often relying on synthetic or social media data, and lack frameworks to assess when automated judges can be trusted. To address the need for large-scale dialogue datasets and judge reliability assessment, we introduce two benchmarks that provide a framework for generation and evaluation. MentalBench-100k consolidates 10,000 one-turn conversations from three real scenarios datasets, each paired with nine LLM-generated responses, yielding 100,000 response pairs. MentalAlign-70k}reframes evaluation by comparing four high-performing LLM judges with human experts across 70,000 ratings on seven attributes, grouped into Cognitive Support Score (CSS) and Affective Resonance Score (ARS). We then employ the Affective Cognitive Agreement Framework, a statistical methodology using intraclass correlation coefficients (ICC) with confidence intervals to quantify agreement, consistency, and bias between LLM judges and human experts. Our analysis reveals systematic inflation by LLM judges, strong reliability for cognitive attributes such as guidance and informativeness, reduced precision for empathy, and some unreliability in safety and relevance. Our contributions establish new methodological and empirical foundations for reliable, large-scale evaluation of LLMs in mental health. We release the benchmarks and codes at: https://github.com/abeerbadawi/MentalBench/
format Preprint
id arxiv_https___arxiv_org_abs_2510_19032
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
Badawi, Abeer
Rahimi, Elahe
Laskar, Md Tahmid Rahman
Grach, Sheri
Bertrand, Lindsay
Danok, Lames
Huang, Jimmy
Rudzicz, Frank
Dolatabadi, Elham
Computation and Language
Computers and Society
Human-Computer Interaction
Evaluating Large Language Models (LLMs) for mental health support is challenging due to the emotionally and cognitively complex nature of therapeutic dialogue. Existing benchmarks are limited in scale, reliability, often relying on synthetic or social media data, and lack frameworks to assess when automated judges can be trusted. To address the need for large-scale dialogue datasets and judge reliability assessment, we introduce two benchmarks that provide a framework for generation and evaluation. MentalBench-100k consolidates 10,000 one-turn conversations from three real scenarios datasets, each paired with nine LLM-generated responses, yielding 100,000 response pairs. MentalAlign-70k}reframes evaluation by comparing four high-performing LLM judges with human experts across 70,000 ratings on seven attributes, grouped into Cognitive Support Score (CSS) and Affective Resonance Score (ARS). We then employ the Affective Cognitive Agreement Framework, a statistical methodology using intraclass correlation coefficients (ICC) with confidence intervals to quantify agreement, consistency, and bias between LLM judges and human experts. Our analysis reveals systematic inflation by LLM judges, strong reliability for cognitive attributes such as guidance and informativeness, reduced precision for empathy, and some unreliability in safety and relevance. Our contributions establish new methodological and empirical foundations for reliable, large-scale evaluation of LLMs in mental health. We release the benchmarks and codes at: https://github.com/abeerbadawi/MentalBench/
title When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
topic Computation and Language
Computers and Society
Human-Computer Interaction
url https://arxiv.org/abs/2510.19032