TherapyGym: Evaluating and Aligning Clinical Fidelity and Safety in Therapy Chatbots
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Fangrui, Chbeir, Souhad, Khatua, Arpandeep, Wang, Sheng, Tan, Sijun, Ye, Kenan, Bailey, Lily, Daniel, Merryn, Louie, Ryan, Koyejo, Sanmi, Adeli, Ehsan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
di: Vo, Truong, et al.
Pubblicazione: (2025)
di: Vo, Truong, et al.
Pubblicazione: (2025)
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
di: Durante, Zane, et al.
Pubblicazione: (2026)
di: Durante, Zane, et al.
Pubblicazione: (2026)
Why Do Safety Guardrails Degrade Across Languages?
di: Zhang, Max, et al.
Pubblicazione: (2026)
di: Zhang, Max, et al.
Pubblicazione: (2026)
Cycle Diffusion Model for Counterfactual Image Generation
di: Huang, Fangrui, et al.
Pubblicazione: (2025)
di: Huang, Fangrui, et al.
Pubblicazione: (2025)
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
di: Kugo, Noriyuki, et al.
Pubblicazione: (2025)
di: Kugo, Noriyuki, et al.
Pubblicazione: (2025)
Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models
di: Semnani, Sina J., et al.
Pubblicazione: (2025)
di: Semnani, Sina J., et al.
Pubblicazione: (2025)
Discovering Implicit Large Language Model Alignment Objectives
di: Chen, Edward, et al.
Pubblicazione: (2026)
di: Chen, Edward, et al.
Pubblicazione: (2026)
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
di: Wang, Angelina, et al.
Pubblicazione: (2025)
di: Wang, Angelina, et al.
Pubblicazione: (2025)
TherapyProbe: Generating Design Knowledge for Relational Safety in Mental Health Chatbots Through Adversarial Simulation
di: Chandra, Joydeep, et al.
Pubblicazione: (2026)
di: Chandra, Joydeep, et al.
Pubblicazione: (2026)
A Gentle Approach to Multi-Sensor Fusion Data Using Linear Kalman Filter
di: Veysi, Parsa, et al.
Pubblicazione: (2024)
di: Veysi, Parsa, et al.
Pubblicazione: (2024)
Reasoning Models Don't Just Think Longer, They Move Differently
di: Gjølbye, Anders, et al.
Pubblicazione: (2026)
di: Gjølbye, Anders, et al.
Pubblicazione: (2026)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
di: Wang, Angelina, et al.
Pubblicazione: (2025)
di: Wang, Angelina, et al.
Pubblicazione: (2025)
HiFA: High-fidelity Text-to-3D Generation with Advanced Diffusion Guidance
di: Zhu, Junzhe, et al.
Pubblicazione: (2023)
di: Zhu, Junzhe, et al.
Pubblicazione: (2023)
Discovering Latent Graphs with GFlowNets for Diverse Conditional Image Generation
di: Trang, Bailey, et al.
Pubblicazione: (2025)
di: Trang, Bailey, et al.
Pubblicazione: (2025)
Transforming and Combining Rewards for Aligning Large Language Models
di: Wang, Zihao, et al.
Pubblicazione: (2024)
di: Wang, Zihao, et al.
Pubblicazione: (2024)
Logits are All We Need to Adapt Closed Models
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
Curriculum-Guided Layer Scaling for Language Model Pretraining
di: Singh, Karanpartap, et al.
Pubblicazione: (2025)
di: Singh, Karanpartap, et al.
Pubblicazione: (2025)
SpecEval: Evaluating Model Adherence to Behavior Specifications
di: Ahmed, Ahmed, et al.
Pubblicazione: (2025)
di: Ahmed, Ahmed, et al.
Pubblicazione: (2025)
Do You Understand How I Feel?: Towards Verified Empathy in Therapy Chatbots
di: Dettori, Francesco, et al.
Pubblicazione: (2026)
di: Dettori, Francesco, et al.
Pubblicazione: (2026)
Position: Model Collapse Does Not Mean What You Think
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2025)
Extracting books from production language models
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026)
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026)
Is Pre-training Truly Better Than Meta-Learning?
di: Miranda, Brando, et al.
Pubblicazione: (2023)
di: Miranda, Brando, et al.
Pubblicazione: (2023)
Scalable Ensembling For Mitigating Reward Overoptimisation
di: Ahmed, Ahmed M., et al.
Pubblicazione: (2024)
di: Ahmed, Ahmed M., et al.
Pubblicazione: (2024)
Reliable and Efficient Amortized Model-based Evaluation
di: Truong, Sang, et al.
Pubblicazione: (2025)
di: Truong, Sang, et al.
Pubblicazione: (2025)
Differentially Private Adaptation of Diffusion Models via Noisy Aggregated Embeddings
di: Peetathawatchai, Pura, et al.
Pubblicazione: (2024)
di: Peetathawatchai, Pura, et al.
Pubblicazione: (2024)
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
di: Haupt, Andreas, et al.
Pubblicazione: (2026)
di: Haupt, Andreas, et al.
Pubblicazione: (2026)
How Real Are Synthetic Therapy Conversations? Evaluating Fidelity in Prolonged Exposure Dialogues
di: BN, Suhas, et al.
Pubblicazione: (2025)
di: BN, Suhas, et al.
Pubblicazione: (2025)
AdaVid: Adaptive Video-Language Pretraining
di: Patel, Chaitanya, et al.
Pubblicazione: (2025)
di: Patel, Chaitanya, et al.
Pubblicazione: (2025)
On Fairness of Low-Rank Adaptation of Large Models
di: Ding, Zhoujie, et al.
Pubblicazione: (2024)
di: Ding, Zhoujie, et al.
Pubblicazione: (2024)
Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs
di: Panda, Ashwinee, et al.
Pubblicazione: (2024)
di: Panda, Ashwinee, et al.
Pubblicazione: (2024)
Scaling Laws for Downstream Task Performance of Large Language Models
di: Isik, Berivan, et al.
Pubblicazione: (2024)
di: Isik, Berivan, et al.
Pubblicazione: (2024)
Artist-Created Mesh Generation from Raw Observation
di: He, Yao, et al.
Pubblicazione: (2025)
di: He, Yao, et al.
Pubblicazione: (2025)
Towards Robust 3D Pose Transfer with Adversarial Learning
di: Chen, Haoyu, et al.
Pubblicazione: (2024)
di: Chen, Haoyu, et al.
Pubblicazione: (2024)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
di: Miranda, Brando, et al.
Pubblicazione: (2023)
di: Miranda, Brando, et al.
Pubblicazione: (2023)
STONet: A neural operator for modeling solute transport in micro-cracked reservoirs
di: Haghighat, Ehsan, et al.
Pubblicazione: (2024)
di: Haghighat, Ehsan, et al.
Pubblicazione: (2024)
The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models
di: Paruchuri, Akshay, et al.
Pubblicazione: (2026)
di: Paruchuri, Akshay, et al.
Pubblicazione: (2026)
Steering Away from Memorization: Reachability-Constrained Reinforcement Learning for Text-to-Image Diffusion
di: Karnik, Sathwik, et al.
Pubblicazione: (2026)
di: Karnik, Sathwik, et al.
Pubblicazione: (2026)
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
di: Iyer, Laya, et al.
Pubblicazione: (2026)
di: Iyer, Laya, et al.
Pubblicazione: (2026)
NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding
di: Abbasi, Mohammad H., et al.
Pubblicazione: (2026)
di: Abbasi, Mohammad H., et al.
Pubblicazione: (2026)
Documenti analoghi
-
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
di: Vo, Truong, et al.
Pubblicazione: (2025) -
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
di: Durante, Zane, et al.
Pubblicazione: (2026) -
Why Do Safety Guardrails Degrade Across Languages?
di: Zhang, Max, et al.
Pubblicazione: (2026) -
Cycle Diffusion Model for Counterfactual Image Generation
di: Huang, Fangrui, et al.
Pubblicazione: (2025) -
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
di: Kugo, Noriyuki, et al.
Pubblicazione: (2025)