BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Yubin, Hu, Zhiyuan, Jeong, Hyewon, Park, Eugene, Li, Shuyue Stella, Park, Chanwoo, Xiong, Shiyun, Lu, MingYu, Lee, Hyeonhoon, Liu, Xin, McDuff, Daniel, Breazeal, Cynthia, Tulebaev, Samir, Park, Hae Won
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912399237316608
author Kim, Yubin
Hu, Zhiyuan
Jeong, Hyewon
Park, Eugene
Li, Shuyue Stella
Park, Chanwoo
Xiong, Shiyun
Lu, MingYu
Lee, Hyeonhoon
Liu, Xin
McDuff, Daniel
Breazeal, Cynthia
Tulebaev, Samir
Park, Hae Won
author_facet Kim, Yubin
Hu, Zhiyuan
Jeong, Hyewon
Park, Eugene
Li, Shuyue Stella
Park, Chanwoo
Xiong, Shiyun
Lu, MingYu
Lee, Hyeonhoon
Liu, Xin
McDuff, Daniel
Breazeal, Cynthia
Tulebaev, Samir
Park, Hae Won
contents Large Language Models (LLMs) as clinical agents require careful behavioral adaptation. While adept at reactive tasks (e.g., diagnosis reasoning), LLMs often struggle with proactive engagement, like unprompted identification of critical missing information or risks. We introduce BehaviorBench, a comprehensive dataset to evaluate agent behaviors across a clinical assistance spectrum, ranging from reactive query responses to proactive interventions (e.g., clarifying ambiguities, flagging overlooked critical data). Our BehaviorBench experiments reveal LLMs' inconsistent proactivity. To address this, we propose BehaviorSFT, a novel training strategy using behavioral tokens to explicitly condition LLMs for dynamic behavioral selection along this spectrum. BehaviorSFT boosts performance, achieving up to 97.3% overall Macro F1 on BehaviorBench and improving proactive task scores (e.g., from 95.0% to 96.5% for Qwen2.5-7B-Ins). Crucially, blind clinician evaluations confirmed BehaviorSFT-trained agents exhibit more realistic clinical behavior, striking a superior balance between helpful proactivity (e.g., timely, relevant suggestions) and necessary restraint (e.g., avoiding over-intervention) versus standard fine-tuning or explicit instructed agents.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21757
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum
Kim, Yubin
Hu, Zhiyuan
Jeong, Hyewon
Park, Eugene
Li, Shuyue Stella
Park, Chanwoo
Xiong, Shiyun
Lu, MingYu
Lee, Hyeonhoon
Liu, Xin
McDuff, Daniel
Breazeal, Cynthia
Tulebaev, Samir
Park, Hae Won
Computation and Language
Large Language Models (LLMs) as clinical agents require careful behavioral adaptation. While adept at reactive tasks (e.g., diagnosis reasoning), LLMs often struggle with proactive engagement, like unprompted identification of critical missing information or risks. We introduce BehaviorBench, a comprehensive dataset to evaluate agent behaviors across a clinical assistance spectrum, ranging from reactive query responses to proactive interventions (e.g., clarifying ambiguities, flagging overlooked critical data). Our BehaviorBench experiments reveal LLMs' inconsistent proactivity. To address this, we propose BehaviorSFT, a novel training strategy using behavioral tokens to explicitly condition LLMs for dynamic behavioral selection along this spectrum. BehaviorSFT boosts performance, achieving up to 97.3% overall Macro F1 on BehaviorBench and improving proactive task scores (e.g., from 95.0% to 96.5% for Qwen2.5-7B-Ins). Crucially, blind clinician evaluations confirmed BehaviorSFT-trained agents exhibit more realistic clinical behavior, striking a superior balance between helpful proactivity (e.g., timely, relevant suggestions) and necessary restraint (e.g., avoiding over-intervention) versus standard fine-tuning or explicit instructed agents.
title BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum
topic Computation and Language
url https://arxiv.org/abs/2505.21757