Mitigating Harmful Erraticism in LLMs Through Dialectical Behavior Therapy Based De-Escalation Strategies
Fuente:
arXiv
Saved in:
| Main Authors: | Rangarajan, Pooja, Boyle, Jacob |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Alignment Drift in Multimodal LLMs: A Two-Phase, Longitudinal Evaluation of Harm Across Eight Model Releases
by: Ford, Casey, et al.
Published: (2026)
by: Ford, Casey, et al.
Published: (2026)
Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems
by: Cheng, Myra, et al.
Published: (2025)
by: Cheng, Myra, et al.
Published: (2025)
Getting out of the Big-Muddy: Escalation of Commitment in LLMs
by: Barkett, Emilio, et al.
Published: (2025)
by: Barkett, Emilio, et al.
Published: (2025)
Strategic Chain-of-Thought: Guiding Accurate Reasoning in LLMs through Strategy Elicitation
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
Detecting and Preventing Harmful Behaviors in AI Companions: Development and Evaluation of the SHIELD Supervisory System
by: Ben-Zion, Ziv, et al.
Published: (2025)
by: Ben-Zion, Ziv, et al.
Published: (2025)
TherapyProbe: Generating Design Knowledge for Relational Safety in Mental Health Chatbots Through Adversarial Simulation
by: Chandra, Joydeep, et al.
Published: (2026)
by: Chandra, Joydeep, et al.
Published: (2026)
Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
An AI-Powered Research Assistant in the Lab: A Practical Guide for Text Analysis Through Iterative Collaboration with LLMs
by: Carmona-Díaz, Gino, et al.
Published: (2025)
by: Carmona-Díaz, Gino, et al.
Published: (2025)
Observing Dialogue in Therapy: Categorizing and Forecasting Behavioral Codes
by: Cao, Jie, et al.
Published: (2019)
by: Cao, Jie, et al.
Published: (2019)
The Emotional Spectrum of LLMs: Leveraging Empathy and Emotion-Based Markers for Mental Health Support
by: De Grandi, Alessandro, et al.
Published: (2024)
by: De Grandi, Alessandro, et al.
Published: (2024)
TheraGen: Therapy for Every Generation
by: Doshi, Kartikey, et al.
Published: (2024)
by: Doshi, Kartikey, et al.
Published: (2024)
Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
by: McCullum, Lucas, et al.
Published: (2025)
by: McCullum, Lucas, et al.
Published: (2025)
Adjust for Trust: Mitigating Trust-Induced Inappropriate Reliance on AI Assistance
by: Srinivasan, Tejas, et al.
Published: (2025)
by: Srinivasan, Tejas, et al.
Published: (2025)
NLPGuard: A Framework for Mitigating the Use of Protected Attributes by NLP Classifiers
by: Greco, Salvatore, et al.
Published: (2024)
by: Greco, Salvatore, et al.
Published: (2024)
User-Assistant Bias in LLMs
by: Pan, Xu, et al.
Published: (2025)
by: Pan, Xu, et al.
Published: (2025)
HICode: Hierarchical Inductive Coding with LLMs
by: Zhong, Mian, et al.
Published: (2025)
by: Zhong, Mian, et al.
Published: (2025)
Harmful Traits of AI Companions
by: Knox, W. Bradley, et al.
Published: (2025)
by: Knox, W. Bradley, et al.
Published: (2025)
Effects of Varying LLM Access on Essay Writing Behavior
by: Christenson, Julia, et al.
Published: (2026)
by: Christenson, Julia, et al.
Published: (2026)
Offscript: Automated Auditing of Instruction Adherence in LLMs
by: Clark, Nicholas, et al.
Published: (2025)
by: Clark, Nicholas, et al.
Published: (2025)
Can LLMs Generate Visualizations with Dataless Prompts?
by: Coelho, Darius, et al.
Published: (2024)
by: Coelho, Darius, et al.
Published: (2024)
Evaluating LLMs as Human Surrogates in Controlled Experiments
by: Hoq, Adnan, et al.
Published: (2026)
by: Hoq, Adnan, et al.
Published: (2026)
Aligning LLMs with Individual Preferences via Interaction
by: Wu, Shujin, et al.
Published: (2024)
by: Wu, Shujin, et al.
Published: (2024)
Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments
by: Daynauth, Roland, et al.
Published: (2024)
by: Daynauth, Roland, et al.
Published: (2024)
Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions
by: Fan, Xianzhe, et al.
Published: (2024)
by: Fan, Xianzhe, et al.
Published: (2024)
An Analysis of User Behaviors for Objectively Evaluating Spoken Dialogue Systems
by: Inoue, Koji, et al.
Published: (2024)
by: Inoue, Koji, et al.
Published: (2024)
Are Today's LLMs Ready to Explain Well-Being Concepts?
by: Jiang, Bohan, et al.
Published: (2025)
by: Jiang, Bohan, et al.
Published: (2025)
Clinical knowledge in LLMs does not translate to human interactions
by: Bean, Andrew M., et al.
Published: (2025)
by: Bean, Andrew M., et al.
Published: (2025)
AutoLife: Automatic Life Journaling with Smartphones and LLMs
by: Xu, Huatao, et al.
Published: (2024)
by: Xu, Huatao, et al.
Published: (2024)
HARGPT: Are LLMs Zero-Shot Human Activity Recognizers?
by: Ji, Sijie, et al.
Published: (2024)
by: Ji, Sijie, et al.
Published: (2024)
Can Large Language Model Agents Simulate Human Trust Behavior?
by: Xie, Chengxing, et al.
Published: (2024)
by: Xie, Chengxing, et al.
Published: (2024)
Enhancing Public Speaking Skills in Engineering Students Through AI
by: Harsh, Amol, et al.
Published: (2025)
by: Harsh, Amol, et al.
Published: (2025)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
Direct Advantage Regression: Aligning LLMs with Online AI Reward
by: He, Li, et al.
Published: (2025)
by: He, Li, et al.
Published: (2025)
Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs
by: Zhu, Shaojie, et al.
Published: (2023)
by: Zhu, Shaojie, et al.
Published: (2023)
AI Conversational Interviewing: Transforming Surveys with LLMs as Adaptive Interviewers
by: Wuttke, Alexander, et al.
Published: (2024)
by: Wuttke, Alexander, et al.
Published: (2024)
CBEval: A framework for evaluating and interpreting cognitive biases in LLMs
by: Shaikh, Ammar, et al.
Published: (2024)
by: Shaikh, Ammar, et al.
Published: (2024)
Copiloting Diagnosis of Autism in Real Clinical Scenarios via LLMs
by: Jiang, Yi, et al.
Published: (2024)
by: Jiang, Yi, et al.
Published: (2024)
Exploring Human Perceptions of AI Responses: Insights from a Mixed-Methods Study on Risk Mitigation in Generative Models
by: Candello, Heloisa, et al.
Published: (2025)
by: Candello, Heloisa, et al.
Published: (2025)
Meta-Evaluating Local LLMs: Rethinking Performance Metrics for Serious Games
by: Isaza-Giraldo, Andrés, et al.
Published: (2025)
by: Isaza-Giraldo, Andrés, et al.
Published: (2025)
Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?
by: Shen, Hua, et al.
Published: (2025)
by: Shen, Hua, et al.
Published: (2025)
Similar Items
-
Alignment Drift in Multimodal LLMs: A Two-Phase, Longitudinal Evaluation of Harm Across Eight Model Releases
by: Ford, Casey, et al.
Published: (2026) -
Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems
by: Cheng, Myra, et al.
Published: (2025) -
Getting out of the Big-Muddy: Escalation of Commitment in LLMs
by: Barkett, Emilio, et al.
Published: (2025) -
Strategic Chain-of-Thought: Guiding Accurate Reasoning in LLMs through Strategy Elicitation
by: Wang, Yu, et al.
Published: (2024) -
Detecting and Preventing Harmful Behaviors in AI Companions: Development and Evaluation of the SHIELD Supervisory System
by: Ben-Zion, Ziv, et al.
Published: (2025)