PSI-Bench: Towards Clinically Grounded and Interpretable Evaluation of Depression Patient Simulators
Fuente:
arXiv
Salvato in:
| Autori principali: | Hoang, Nguyen Khoi, Mehri, Shuhaib, Hsu, Tse-An, Sun, Yi-Jyun, Truong, Quynh Xuan Nguyen, Doan, Khoa D, Hakkani-Tür, Dilek |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Goal Alignment in LLM-Based User Simulators for Conversational AI
di: Mehri, Shuhaib, et al.
Pubblicazione: (2025)
di: Mehri, Shuhaib, et al.
Pubblicazione: (2025)
Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models
di: Bozdag, Nimet Beyza, et al.
Pubblicazione: (2025)
di: Bozdag, Nimet Beyza, et al.
Pubblicazione: (2025)
Beyond Sample-Level Feedback: Using Reference-Level Feedback to Guide Data Synthesis
di: Mehri, Shuhaib, et al.
Pubblicazione: (2025)
di: Mehri, Shuhaib, et al.
Pubblicazione: (2025)
Sparking Scientific Creativity via LLM-Driven Interdisciplinary Inspiration
di: Kargupta, Priyanka, et al.
Pubblicazione: (2026)
di: Kargupta, Priyanka, et al.
Pubblicazione: (2026)
MultiSessionCollab: Learning User Preferences with Memory to Improve Long-Term Collaboration
di: Mehri, Shuhaib, et al.
Pubblicazione: (2026)
di: Mehri, Shuhaib, et al.
Pubblicazione: (2026)
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction
di: Hao, Yuren, et al.
Pubblicazione: (2026)
di: Hao, Yuren, et al.
Pubblicazione: (2026)
Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
di: Mehri, Shuhaib, et al.
Pubblicazione: (2026)
di: Mehri, Shuhaib, et al.
Pubblicazione: (2026)
Know Your Mistakes: Towards Preventing Overreliance on Task-Oriented Conversational AI Through Accountability Modeling
di: Dey, Suvodip, et al.
Pubblicazione: (2025)
di: Dey, Suvodip, et al.
Pubblicazione: (2025)
Confidence Estimation for LLM-Based Dialogue State Tracking
di: Sun, Yi-Jyun, et al.
Pubblicazione: (2024)
di: Sun, Yi-Jyun, et al.
Pubblicazione: (2024)
Simulating User Agents for Embodied Conversational-AI
di: Philipov, Daniel, et al.
Pubblicazione: (2024)
di: Philipov, Daniel, et al.
Pubblicazione: (2024)
AURA: A Diagnostic Framework for Tracking User Satisfaction of Interactive Planning Agents
di: Kim, Takyoung, et al.
Pubblicazione: (2025)
di: Kim, Takyoung, et al.
Pubblicazione: (2025)
Do LLMs Encode Functional Importance of Reasoning Tokens?
di: Singh, Janvijay, et al.
Pubblicazione: (2026)
di: Singh, Janvijay, et al.
Pubblicazione: (2026)
Neural Networks for Learnable and Scalable Influence Estimation of Instruction Fine-Tuning Data
di: Agarwal, Ishika, et al.
Pubblicazione: (2025)
di: Agarwal, Ishika, et al.
Pubblicazione: (2025)
Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue
di: Dongre, Vardhan, et al.
Pubblicazione: (2026)
di: Dongre, Vardhan, et al.
Pubblicazione: (2026)
Must Read: A Comprehensive Survey of Computational Persuasion
di: Bozdag, Nimet Beyza, et al.
Pubblicazione: (2025)
di: Bozdag, Nimet Beyza, et al.
Pubblicazione: (2025)
YourBench: Easy Custom Evaluation Sets for Everyone
di: Shashidhar, Sumuk, et al.
Pubblicazione: (2025)
di: Shashidhar, Sumuk, et al.
Pubblicazione: (2025)
Plan Verification for LLM-Based Embodied Task Completion Agents
di: Hariharan, Ananth, et al.
Pubblicazione: (2025)
di: Hariharan, Ananth, et al.
Pubblicazione: (2025)
SIMU: Selective Influence Machine Unlearning
di: Agarwal, Anu, et al.
Pubblicazione: (2025)
di: Agarwal, Anu, et al.
Pubblicazione: (2025)
Question Generation for Assessing Early Literacy Reading Comprehension
di: Yang, Xiaocheng, et al.
Pubblicazione: (2025)
di: Yang, Xiaocheng, et al.
Pubblicazione: (2025)
From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
di: Rabbani, Parisa, et al.
Pubblicazione: (2025)
di: Rabbani, Parisa, et al.
Pubblicazione: (2025)
Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems
di: Kazi, Taaha, et al.
Pubblicazione: (2024)
di: Kazi, Taaha, et al.
Pubblicazione: (2024)
ReSpAct: Harmonizing Reasoning, Speaking, and Acting Towards Building Large Language Model-Based Conversational AI Agents
di: Dongre, Vardhan, et al.
Pubblicazione: (2024)
di: Dongre, Vardhan, et al.
Pubblicazione: (2024)
Unveiling Concept Attribution in Diffusion Models
di: Nguyen, Quang H., et al.
Pubblicazione: (2024)
di: Nguyen, Quang H., et al.
Pubblicazione: (2024)
Pauli nonlocality and the nucleon effective mass
di: Khoa, Dao T., et al.
Pubblicazione: (2024)
di: Khoa, Dao T., et al.
Pubblicazione: (2024)
Reinforcement Learning Finetunes Small Subnetworks in Large Language Models
di: Mukherjee, Sagnik, et al.
Pubblicazione: (2025)
di: Mukherjee, Sagnik, et al.
Pubblicazione: (2025)
Unsupervised Human Preference Learning
di: Shashidhar, Sumuk, et al.
Pubblicazione: (2024)
di: Shashidhar, Sumuk, et al.
Pubblicazione: (2024)
A Rising Tide Lifts All Boats: MTQE Rewards for Idioms Improve General Translation Quality
di: Agarwal, Ishika, et al.
Pubblicazione: (2026)
di: Agarwal, Ishika, et al.
Pubblicazione: (2026)
Instruct, Not Assist: LLM-based Multi-Turn Planning and Hierarchical Questioning for Socratic Code Debugging
di: Kargupta, Priyanka, et al.
Pubblicazione: (2024)
di: Kargupta, Priyanka, et al.
Pubblicazione: (2024)
LLMs are Vulnerable to Malicious Prompts Disguised as Scientific Language
di: Ge, Yubin, et al.
Pubblicazione: (2025)
di: Ge, Yubin, et al.
Pubblicazione: (2025)
ReasoningFlow: Semantic Structure of Complex Reasoning Traces
di: Lee, Jinu, et al.
Pubblicazione: (2025)
di: Lee, Jinu, et al.
Pubblicazione: (2025)
Self-Improving LLM Agents at Test-Time
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2025)
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2025)
DocCHA: Towards LLM-Augmented Interactive Online diagnosis System
di: Liu, Xinyi, et al.
Pubblicazione: (2025)
di: Liu, Xinyi, et al.
Pubblicazione: (2025)
Multi-view Action Recognition via Directed Gromov-Wasserstein Discrepancy
di: Nguyen, Hoang-Quan, et al.
Pubblicazione: (2024)
di: Nguyen, Hoang-Quan, et al.
Pubblicazione: (2024)
Venomancer: Towards Imperceptible and Target-on-Demand Backdoor Attacks in Federated Learning
di: Nguyen, Son, et al.
Pubblicazione: (2024)
di: Nguyen, Son, et al.
Pubblicazione: (2024)
Language Specific Knowledge: Do Models Know Better in X than in English?
di: Agarwal, Ishika, et al.
Pubblicazione: (2025)
di: Agarwal, Ishika, et al.
Pubblicazione: (2025)
Optimal PSI-Based Load Shedding Strategy for Static Voltage Stability Enhancement in Islanded Microgrids
di: Trong Nghia, Le, et al.
Pubblicazione: (2026)
di: Trong Nghia, Le, et al.
Pubblicazione: (2026)
Interpreting Microbiome Signatures with MicrobiomeNet
di: Yao Lu, et al.
Pubblicazione: (2026)
di: Yao Lu, et al.
Pubblicazione: (2026)
Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2026)
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2026)
LP-OVOD: Open-Vocabulary Object Detection by Linear Probing
di: Pham, Chau, et al.
Pubblicazione: (2023)
di: Pham, Chau, et al.
Pubblicazione: (2023)
Detecting Neurovascular Instability from Multimodal Physiological Signals Using Wearable-Compatible Edge AI: A Responsible Computational Framework
di: Hoa, Truong Quynh, et al.
Pubblicazione: (2026)
di: Hoa, Truong Quynh, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Goal Alignment in LLM-Based User Simulators for Conversational AI
di: Mehri, Shuhaib, et al.
Pubblicazione: (2025) -
Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models
di: Bozdag, Nimet Beyza, et al.
Pubblicazione: (2025) -
Beyond Sample-Level Feedback: Using Reference-Level Feedback to Guide Data Synthesis
di: Mehri, Shuhaib, et al.
Pubblicazione: (2025) -
Sparking Scientific Creativity via LLM-Driven Interdisciplinary Inspiration
di: Kargupta, Priyanka, et al.
Pubblicazione: (2026) -
MultiSessionCollab: Learning User Preferences with Memory to Improve Long-Term Collaboration
di: Mehri, Shuhaib, et al.
Pubblicazione: (2026)