TN-Eval: Rubric and Evaluation Protocols for Measuring the Quality of Behavioral Therapy Notes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shah, Raj Sanjay, Xu, Lei, Liu, Qianchu, Burnsky, Jon, Bertagnolli, Drew, Shivade, Chaitanya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912295419904000
author Shah, Raj Sanjay
Xu, Lei
Liu, Qianchu
Burnsky, Jon
Bertagnolli, Drew
Shivade, Chaitanya
author_facet Shah, Raj Sanjay
Xu, Lei
Liu, Qianchu
Burnsky, Jon
Bertagnolli, Drew
Shivade, Chaitanya
contents Behavioral therapy notes are important for both legal compliance and patient care. Unlike progress notes in physical health, quality standards for behavioral therapy notes remain underdeveloped. To address this gap, we collaborated with licensed therapists to design a comprehensive rubric for evaluating therapy notes across key dimensions: completeness, conciseness, and faithfulness. Further, we extend a public dataset of behavioral health conversations with therapist-written notes and LLM-generated notes, and apply our evaluation framework to measure their quality. We find that: (1) A rubric-based manual evaluation protocol offers more reliable and interpretable results than traditional Likert-scale annotations. (2) LLMs can mimic human evaluators in assessing completeness and conciseness but struggle with faithfulness. (3) Therapist-written notes often lack completeness and conciseness, while LLM-generated notes contain hallucination. Surprisingly, in a blind test, therapists prefer and judge LLM-generated notes to be superior to therapist-written notes.
format Preprint
id arxiv_https___arxiv_org_abs_2503_20648
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TN-Eval: Rubric and Evaluation Protocols for Measuring the Quality of Behavioral Therapy Notes
Shah, Raj Sanjay
Xu, Lei
Liu, Qianchu
Burnsky, Jon
Bertagnolli, Drew
Shivade, Chaitanya
Computation and Language
Artificial Intelligence
Behavioral therapy notes are important for both legal compliance and patient care. Unlike progress notes in physical health, quality standards for behavioral therapy notes remain underdeveloped. To address this gap, we collaborated with licensed therapists to design a comprehensive rubric for evaluating therapy notes across key dimensions: completeness, conciseness, and faithfulness. Further, we extend a public dataset of behavioral health conversations with therapist-written notes and LLM-generated notes, and apply our evaluation framework to measure their quality. We find that: (1) A rubric-based manual evaluation protocol offers more reliable and interpretable results than traditional Likert-scale annotations. (2) LLMs can mimic human evaluators in assessing completeness and conciseness but struggle with faithfulness. (3) Therapist-written notes often lack completeness and conciseness, while LLM-generated notes contain hallucination. Surprisingly, in a blind test, therapists prefer and judge LLM-generated notes to be superior to therapist-written notes.
title TN-Eval: Rubric and Evaluation Protocols for Measuring the Quality of Behavioral Therapy Notes
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.20648