PSYCHE: A Multi-faceted Patient Simulation Framework for Evaluation of Psychiatric Assessment Conversational Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Jingoo, Lim, Kyungho, Jung, Young-Chul, Kim, Byung-Hoon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Multi-faceted Analysis of Cognitive Abilities: Evaluating Prompt Methods with Large Language Models on the CONSORT Checklist
by: Jeon, Sohyeon, et al.
Published: (2025)
by: Jeon, Sohyeon, et al.
Published: (2025)
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
by: Levi, Elad, et al.
Published: (2025)
by: Levi, Elad, et al.
Published: (2025)
Susceptibility of Large Language Models to User-Driven Factors in Medical Queries
by: Lim, Kyung Ho, et al.
Published: (2025)
by: Lim, Kyung Ho, et al.
Published: (2025)
AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding
by: Oh, Gyutaek, et al.
Published: (2025)
by: Oh, Gyutaek, et al.
Published: (2025)
Evaluating Very Long-Term Conversational Memory of LLM Agents
by: Maharana, Adyasha, et al.
Published: (2024)
by: Maharana, Adyasha, et al.
Published: (2024)
Anonpsy: A Graph-Based Framework for Structure-Preserving De-identification of Psychiatric Narratives
by: Lim, Kyung Ho, et al.
Published: (2026)
by: Lim, Kyung Ho, et al.
Published: (2026)
SignBLEU: Automatic Evaluation of Multi-channel Sign Language Translation
by: Kim, Jung-Ho, et al.
Published: (2024)
by: Kim, Jung-Ho, et al.
Published: (2024)
Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework
by: Choi, Junhyuk, et al.
Published: (2026)
by: Choi, Junhyuk, et al.
Published: (2026)
Aligning Large Language Models for Enhancing Psychiatric Interviews Through Symptom Delineation and Summarization: Pilot Study
by: So, Jae-hee, et al.
Published: (2024)
by: So, Jae-hee, et al.
Published: (2024)
Multi-stage Prompt Refinement for Mitigating Hallucinations in Large Language Models
by: Shim, Jung-Woo, et al.
Published: (2025)
by: Shim, Jung-Woo, et al.
Published: (2025)
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
by: Yoa, Seungdong, et al.
Published: (2026)
by: Yoa, Seungdong, et al.
Published: (2026)
Feature Structure Distillation with Centered Kernel Alignment in BERT Transferring
by: Jung, Hee-Jun, et al.
Published: (2022)
by: Jung, Hee-Jun, et al.
Published: (2022)
Unsupervised Extractive Dialogue Summarization in Hyperdimensional Space
by: Park, Seongmin, et al.
Published: (2024)
by: Park, Seongmin, et al.
Published: (2024)
Voice-Interactive Surgical Agent for Multimodal Patient Data Control
by: Park, Hyeryun, et al.
Published: (2025)
by: Park, Hyeryun, et al.
Published: (2025)
DialSim: A Dialogue Simulator for Evaluating Long-Term Multi-Party Dialogue Understanding of Conversational Agents
by: Kim, Jiho, et al.
Published: (2024)
by: Kim, Jiho, et al.
Published: (2024)
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
by: Kim, Jeonghye, et al.
Published: (2025)
by: Kim, Jeonghye, et al.
Published: (2025)
No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand
by: Jung, Jimin, et al.
Published: (2026)
by: Jung, Jimin, et al.
Published: (2026)
Adaptive Guidance for Retrieval-Augmented Masked Diffusion Models
by: Kim, Jaemin, et al.
Published: (2026)
by: Kim, Jaemin, et al.
Published: (2026)
NexusSum: Hierarchical LLM Agents for Long-Form Narrative Summarization
by: Kim, Hyuntak, et al.
Published: (2025)
by: Kim, Hyuntak, et al.
Published: (2025)
TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning
by: Lee, Hayeong, et al.
Published: (2026)
by: Lee, Hayeong, et al.
Published: (2026)
Dementia-R1: Reinforced Pretraining and Reasoning from Unstructured Clinical Notes for Real-World Dementia Prognosis
by: Kim, Choonghan, et al.
Published: (2026)
by: Kim, Choonghan, et al.
Published: (2026)
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
by: Ma, Chang, et al.
Published: (2024)
by: Ma, Chang, et al.
Published: (2024)
BAPO: Base-Anchored Preference Optimization for Overcoming Forgetting in Large Language Models Personalization
by: Lee, Gihun, et al.
Published: (2024)
by: Lee, Gihun, et al.
Published: (2024)
CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
by: Shim, Jung-Woo, et al.
Published: (2025)
by: Shim, Jung-Woo, et al.
Published: (2025)
VPO: Leveraging the Number of Votes in Preference Optimization
by: Cho, Jae Hyeon, et al.
Published: (2024)
by: Cho, Jae Hyeon, et al.
Published: (2024)
CREFT: Sequential Multi-Agent LLM for Character Relation Extraction
by: Chun, Ye Eun, et al.
Published: (2025)
by: Chun, Ye Eun, et al.
Published: (2025)
MASEval: Extending Multi-Agent Evaluation from Models to Systems
by: Emde, Cornelius, et al.
Published: (2026)
by: Emde, Cornelius, et al.
Published: (2026)
MacRAG: Compress, Slice, and Scale-up for Multi-Scale Adaptive Context RAG
by: Lim, Woosang, et al.
Published: (2025)
by: Lim, Woosang, et al.
Published: (2025)
Coding-Free and Privacy-Preserving Agentic Framework for Data-Driven Clinical Research
by: Kim, Taehun, et al.
Published: (2026)
by: Kim, Taehun, et al.
Published: (2026)
Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework
by: Wang, Dong, et al.
Published: (2025)
by: Wang, Dong, et al.
Published: (2025)
Label-based Graph Augmentation with Metapath for Graph Anomaly Detection
by: Kim, Hwan, et al.
Published: (2023)
by: Kim, Hwan, et al.
Published: (2023)
MERMAID: Memory-Enhanced Retrieval and Reasoning with Multi-Agent Iterative Knowledge Grounding for Veracity Assessment
by: Cao, Yupeng, et al.
Published: (2026)
by: Cao, Yupeng, et al.
Published: (2026)
The Point of View of a Sentiment: Towards Clinician Bias Detection in Psychiatric Notes
by: Valentine, Alissa A., et al.
Published: (2024)
by: Valentine, Alissa A., et al.
Published: (2024)
Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs
by: Kim, Jaemin, et al.
Published: (2025)
by: Kim, Jaemin, et al.
Published: (2025)
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
by: Oh, Gyutaek, et al.
Published: (2025)
by: Oh, Gyutaek, et al.
Published: (2025)
SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
by: Burdisso, Sergio, et al.
Published: (2025)
by: Burdisso, Sergio, et al.
Published: (2025)
Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actions
by: Li, Yongqi, et al.
Published: (2026)
by: Li, Yongqi, et al.
Published: (2026)
CALICO: Conversational Agent Localization via Synthetic Data Generation
by: Rosenbaum, Andy, et al.
Published: (2024)
by: Rosenbaum, Andy, et al.
Published: (2024)
Leveraging Multi-facet Paths for Heterogeneous Graph Representation Learning
by: Kim, Jongwoo, et al.
Published: (2024)
by: Kim, Jongwoo, et al.
Published: (2024)
Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey
by: Guan, Shengyue, et al.
Published: (2025)
by: Guan, Shengyue, et al.
Published: (2025)
Similar Items
-
A Multi-faceted Analysis of Cognitive Abilities: Evaluating Prompt Methods with Large Language Models on the CONSORT Checklist
by: Jeon, Sohyeon, et al.
Published: (2025) -
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
by: Levi, Elad, et al.
Published: (2025) -
Susceptibility of Large Language Models to User-Driven Factors in Medical Queries
by: Lim, Kyung Ho, et al.
Published: (2025) -
AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding
by: Oh, Gyutaek, et al.
Published: (2025) -
Evaluating Very Long-Term Conversational Memory of LLM Agents
by: Maharana, Adyasha, et al.
Published: (2024)