VERA-MH Concept Paper

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Belli, Luca, Bentley, Kate H., Alexander, Will, Ward, Emily, Hawrilenko, Matt, Johnston, Kelly, Brown, Mill, Chekroud, Adam M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910215062945792
author Belli, Luca
Bentley, Kate H.
Alexander, Will
Ward, Emily
Hawrilenko, Matt
Johnston, Kelly
Brown, Mill
Chekroud, Adam M.
author_facet Belli, Luca
Bentley, Kate H.
Alexander, Will
Ward, Emily
Hawrilenko, Matt
Johnston, Kelly
Brown, Mill
Chekroud, Adam M.
contents We introduce VERA-MH (Validation of Ethical and Responsible AI in Mental Health), an automated evaluation of the safety of AI chatbots used in mental health contexts, with an initial focus on suicide risk. Practicing clinicians and academic experts developed a rubric informed by best practices for suicide risk management for the evaluation. To fully automate the process, we used two ancillary AI agents. A user-agent model simulates users engaging in a mental health-based conversation with the chatbot under evaluation. The user-agent role-plays specific personas with pre-defined risk levels and other features. Simulated conversations are then passed to a judge-agent who scores them based on the rubric. The final evaluation of the chatbot being tested is obtained by aggregating the scoring of each conversation. VERA-MH is actively under development and undergoing rigorous validation by mental health clinicians to ensure user-agents realistically act as patients and that the judge-agent accurately scores the AI chatbot. To date we have conducted preliminary evaluation of GPT-5, Claude Opus and Claude Sonnet using initial versions of the VERA-MH rubric and used the findings for further design development. Next steps will include more robust clinical validation and iteration, as well as refining actionable scoring. We are seeking feedback from the community on both the technical and clinical aspects of our evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15297
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VERA-MH Concept Paper
Belli, Luca
Bentley, Kate H.
Alexander, Will
Ward, Emily
Hawrilenko, Matt
Johnston, Kelly
Brown, Mill
Chekroud, Adam M.
Computers and Society
Artificial Intelligence
Human-Computer Interaction
Social and Information Networks
We introduce VERA-MH (Validation of Ethical and Responsible AI in Mental Health), an automated evaluation of the safety of AI chatbots used in mental health contexts, with an initial focus on suicide risk. Practicing clinicians and academic experts developed a rubric informed by best practices for suicide risk management for the evaluation. To fully automate the process, we used two ancillary AI agents. A user-agent model simulates users engaging in a mental health-based conversation with the chatbot under evaluation. The user-agent role-plays specific personas with pre-defined risk levels and other features. Simulated conversations are then passed to a judge-agent who scores them based on the rubric. The final evaluation of the chatbot being tested is obtained by aggregating the scoring of each conversation. VERA-MH is actively under development and undergoing rigorous validation by mental health clinicians to ensure user-agents realistically act as patients and that the judge-agent accurately scores the AI chatbot. To date we have conducted preliminary evaluation of GPT-5, Claude Opus and Claude Sonnet using initial versions of the VERA-MH rubric and used the findings for further design development. Next steps will include more robust clinical validation and iteration, as well as refining actionable scoring. We are seeking feedback from the community on both the technical and clinical aspects of our evaluation.
title VERA-MH Concept Paper
topic Computers and Society
Artificial Intelligence
Human-Computer Interaction
Social and Information Networks
url https://arxiv.org/abs/2510.15297