Do Large Language Models Align with Core Mental Health Counseling Competencies?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912249189236736 |
|---|---|
| author | Nguyen, Viet Cuong Taher, Mohammad Hong, Dongwan Possobom, Vinicius Konkolics Gopalakrishnan, Vibha Thirunellayi Raj, Ekta Li, Zihang Soled, Heather J. Birnbaum, Michael L. Kumar, Srijan De Choudhury, Munmun |
| author_facet | Nguyen, Viet Cuong Taher, Mohammad Hong, Dongwan Possobom, Vinicius Konkolics Gopalakrishnan, Vibha Thirunellayi Raj, Ekta Li, Zihang Soled, Heather J. Birnbaum, Michael L. Kumar, Srijan De Choudhury, Munmun |
| contents | The rapid evolution of Large Language Models (LLMs) presents a promising solution to the global shortage of mental health professionals. However, their alignment with essential counseling competencies remains underexplored. We introduce CounselingBench, a novel NCMHCE-based benchmark evaluating 22 general-purpose and medical-finetuned LLMs across five key competencies. While frontier models surpass minimum aptitude thresholds, they fall short of expert-level performance, excelling in Intake, Assessment & Diagnosis but struggling with Core Counseling Attributes and Professional Practice & Ethics. Surprisingly, medical LLMs do not outperform generalist models in accuracy, though they provide slightly better justifications while making more context-related errors. These findings highlight the challenges of developing AI for mental health counseling, particularly in competencies requiring empathy and nuanced reasoning. Our results underscore the need for specialized, fine-tuned models aligned with core mental health counseling competencies and supported by human oversight before real-world deployment. Code and data associated with this manuscript can be found at: https://github.com/cuongnguyenx/CounselingBench |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_22446 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Do Large Language Models Align with Core Mental Health Counseling Competencies? Nguyen, Viet Cuong Taher, Mohammad Hong, Dongwan Possobom, Vinicius Konkolics Gopalakrishnan, Vibha Thirunellayi Raj, Ekta Li, Zihang Soled, Heather J. Birnbaum, Michael L. Kumar, Srijan De Choudhury, Munmun Computation and Language Artificial Intelligence The rapid evolution of Large Language Models (LLMs) presents a promising solution to the global shortage of mental health professionals. However, their alignment with essential counseling competencies remains underexplored. We introduce CounselingBench, a novel NCMHCE-based benchmark evaluating 22 general-purpose and medical-finetuned LLMs across five key competencies. While frontier models surpass minimum aptitude thresholds, they fall short of expert-level performance, excelling in Intake, Assessment & Diagnosis but struggling with Core Counseling Attributes and Professional Practice & Ethics. Surprisingly, medical LLMs do not outperform generalist models in accuracy, though they provide slightly better justifications while making more context-related errors. These findings highlight the challenges of developing AI for mental health counseling, particularly in competencies requiring empathy and nuanced reasoning. Our results underscore the need for specialized, fine-tuned models aligned with core mental health counseling competencies and supported by human oversight before real-world deployment. Code and data associated with this manuscript can be found at: https://github.com/cuongnguyenx/CounselingBench |
| title | Do Large Language Models Align with Core Mental Health Counseling Competencies? |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2410.22446 |