CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ghosh, Akash, Sridhar, Srivarshinee, Ravi, Raghav Kaushik, Muhsin, Muhsin, Saha, Sriparna, Agarwal, Chirag
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914197450784768
author Ghosh, Akash
Sridhar, Srivarshinee
Ravi, Raghav Kaushik
Muhsin, Muhsin
Saha, Sriparna
Agarwal, Chirag
author_facet Ghosh, Akash
Sridhar, Srivarshinee
Ravi, Raghav Kaushik
Muhsin, Muhsin
Saha, Sriparna
Agarwal, Chirag
contents Integrating language models (LMs) in healthcare systems holds great promise for improving medical workflows and decision-making. However, a critical barrier to their real-world adoption is the lack of reliable evaluation of their trustworthiness, especially in multilingual healthcare settings. Existing LMs are predominantly trained in high-resource languages, making them ill-equipped to handle the complexity and diversity of healthcare queries in mid- and low-resource languages, posing significant challenges for deploying them in global healthcare contexts where linguistic diversity is key. In this work, we present CLINIC, a Comprehensive Multilingual Benchmark to evaluate the trustworthiness of language models in healthcare. CLINIC systematically benchmarks LMs across five key dimensions of trustworthiness: truthfulness, fairness, safety, robustness, and privacy, operationalized through 18 diverse tasks, spanning 15 languages (covering all the major continents), and encompassing a wide array of critical healthcare topics like disease conditions, preventive actions, diagnostic tests, treatments, surgeries, and medications. Our extensive evaluation reveals that LMs struggle with factual correctness, demonstrate bias across demographic and linguistic groups, and are susceptible to privacy breaches and adversarial attacks. By highlighting these shortcomings, CLINIC lays the foundation for enhancing the global reach and safety of LMs in healthcare across diverse languages.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11437
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare
Ghosh, Akash
Sridhar, Srivarshinee
Ravi, Raghav Kaushik
Muhsin, Muhsin
Saha, Sriparna
Agarwal, Chirag
Computation and Language
Integrating language models (LMs) in healthcare systems holds great promise for improving medical workflows and decision-making. However, a critical barrier to their real-world adoption is the lack of reliable evaluation of their trustworthiness, especially in multilingual healthcare settings. Existing LMs are predominantly trained in high-resource languages, making them ill-equipped to handle the complexity and diversity of healthcare queries in mid- and low-resource languages, posing significant challenges for deploying them in global healthcare contexts where linguistic diversity is key. In this work, we present CLINIC, a Comprehensive Multilingual Benchmark to evaluate the trustworthiness of language models in healthcare. CLINIC systematically benchmarks LMs across five key dimensions of trustworthiness: truthfulness, fairness, safety, robustness, and privacy, operationalized through 18 diverse tasks, spanning 15 languages (covering all the major continents), and encompassing a wide array of critical healthcare topics like disease conditions, preventive actions, diagnostic tests, treatments, surgeries, and medications. Our extensive evaluation reveals that LMs struggle with factual correctness, demonstrate bias across demographic and linguistic groups, and are susceptible to privacy breaches and adversarial attacks. By highlighting these shortcomings, CLINIC lays the foundation for enhancing the global reach and safety of LMs in healthcare across diverse languages.
title CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare
topic Computation and Language
url https://arxiv.org/abs/2512.11437