Large Language Models for Mental Health: A Multilingual Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Raihan, Nishat, Puspo, Sadiya Sayara Chowdhury, Bucur, Ana-Maria, Chancellor, Stevie, Zampieri, Marcos
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914301755785216
author Raihan, Nishat
Puspo, Sadiya Sayara Chowdhury
Bucur, Ana-Maria
Chancellor, Stevie
Zampieri, Marcos
author_facet Raihan, Nishat
Puspo, Sadiya Sayara Chowdhury
Bucur, Ana-Maria
Chancellor, Stevie
Zampieri, Marcos
contents Large Language Models (LLMs) have remarkable capabilities across NLP tasks. However, their performance in multilingual contexts, especially within the mental health domain, has not been thoroughly explored. In this paper, we evaluate proprietary and open-source LLMs on eight mental health datasets in various languages, as well as their machine-translated (MT) counterparts. We compare LLM performance in zero-shot, few-shot, and fine-tuned settings against conventional NLP baselines that do not employ LLMs. In addition, we assess translation quality across language families and typologies to understand its influence on LLM performance. Proprietary LLMs and fine-tuned open-source LLMs achieve competitive F1 scores on several datasets, often surpassing state-of-the-art results. However, performance on MT data is generally lower, and the extent of this decline varies by language and typology. This variation highlights both the strengths of LLMs in handling mental health tasks in languages other than English and their limitations when translation quality introduces structural or lexical mismatches.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02440
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Large Language Models for Mental Health: A Multilingual Evaluation
Raihan, Nishat
Puspo, Sadiya Sayara Chowdhury
Bucur, Ana-Maria
Chancellor, Stevie
Zampieri, Marcos
Computation and Language
Large Language Models (LLMs) have remarkable capabilities across NLP tasks. However, their performance in multilingual contexts, especially within the mental health domain, has not been thoroughly explored. In this paper, we evaluate proprietary and open-source LLMs on eight mental health datasets in various languages, as well as their machine-translated (MT) counterparts. We compare LLM performance in zero-shot, few-shot, and fine-tuned settings against conventional NLP baselines that do not employ LLMs. In addition, we assess translation quality across language families and typologies to understand its influence on LLM performance. Proprietary LLMs and fine-tuned open-source LLMs achieve competitive F1 scores on several datasets, often surpassing state-of-the-art results. However, performance on MT data is generally lower, and the extent of this decline varies by language and typology. This variation highlights both the strengths of LLMs in handling mental health tasks in languages other than English and their limitations when translation quality introduces structural or lexical mismatches.
title Large Language Models for Mental Health: A Multilingual Evaluation
topic Computation and Language
url https://arxiv.org/abs/2602.02440