Detecting and Preventing Harmful Behaviors in AI Companions: Development and Evaluation of the SHIELD Supervisory System
Fuente:
arXiv
Saved in:
| Main Authors: | Ben-Zion, Ziv, Raffelhüschen, Paul, Zettl, Max, Lüönd, Antonia, Burrer, Achim, Homan, Philipp, Spiller, Tobias R |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making
by: Ben-Zion, Ziv, et al.
Published: (2025)
by: Ben-Zion, Ziv, et al.
Published: (2025)
Harmful Traits of AI Companions
by: Knox, W. Bradley, et al.
Published: (2025)
by: Knox, W. Bradley, et al.
Published: (2025)
Intimacy as Service, Harm as Externality: Critical Perspectives on AI Companion Platform Accountability
by: Eom, Dayeon, et al.
Published: (2026)
by: Eom, Dayeon, et al.
Published: (2026)
Measuring Machine Companionship: Scale Development and Validation for AI Companions
by: Banks, Jaime
Published: (2025)
by: Banks, Jaime
Published: (2025)
Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions
by: Fan, Xianzhe, et al.
Published: (2024)
by: Fan, Xianzhe, et al.
Published: (2024)
The Rise of AI Companions: Interaction with AI Companions and Psychological Well-being
by: Zhang, Yutong, et al.
Published: (2025)
by: Zhang, Yutong, et al.
Published: (2025)
The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI Relationships
by: Zhang, Renwen, et al.
Published: (2024)
by: Zhang, Renwen, et al.
Published: (2024)
Frictionless Love: Associations Between AI Companion Roles and Behavioral Addiction
by: Agarwal, Vibhor, et al.
Published: (2026)
by: Agarwal, Vibhor, et al.
Published: (2026)
Computational Analysis of Speech Clarity Predicts Audience Engagement in TED Talks
by: Segal, Roni, et al.
Published: (2026)
by: Segal, Roni, et al.
Published: (2026)
Building AI Companions that Prioritise Learning over Performance
by: Khosravi, Hassan, et al.
Published: (2026)
by: Khosravi, Hassan, et al.
Published: (2026)
AI Mismatches: Identifying Potential Algorithmic Harms Before AI Development
by: Saxena, Devansh, et al.
Published: (2025)
by: Saxena, Devansh, et al.
Published: (2025)
Examining Risks in the AI Companion Application Ecosystem
by: Brigham, Natalie Grace, et al.
Published: (2026)
by: Brigham, Natalie Grace, et al.
Published: (2026)
Emotional Manipulation by AI Companions
by: De Freitas, Julian, et al.
Published: (2025)
by: De Freitas, Julian, et al.
Published: (2025)
Surveys Considered Harmful? Reflecting on the Use of Surveys in AI Research, Development, and Governance
by: Tahaei, Mohammmad, et al.
Published: (2024)
by: Tahaei, Mohammmad, et al.
Published: (2024)
Exploring Proactive Interventions toward Harmful Behavior in Embodied Virtual Spaces
by: Panchanadikar, Ruchi
Published: (2024)
by: Panchanadikar, Ruchi
Published: (2024)
Principles of Safe AI Companions for Youth: Parent and Expert Perspectives
by: Yu, Yaman, et al.
Published: (2025)
by: Yu, Yaman, et al.
Published: (2025)
The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers
by: Prather, James, et al.
Published: (2024)
by: Prather, James, et al.
Published: (2024)
Digital Companionship: Overlapping Uses of AI Companions and AI Assistants
by: Manoli, Aikaterina, et al.
Published: (2025)
by: Manoli, Aikaterina, et al.
Published: (2025)
AI Meets the Classroom: When Do Large Language Models Harm Learning?
by: Lehmann, Matthias, et al.
Published: (2024)
by: Lehmann, Matthias, et al.
Published: (2024)
Deletion Considered Harmful
by: Englefield, Paul, et al.
Published: (2025)
by: Englefield, Paul, et al.
Published: (2025)
Generative AI and Perceptual Harms: Who's Suspected of using LLMs?
by: Kadoma, Kowe, et al.
Published: (2024)
by: Kadoma, Kowe, et al.
Published: (2024)
Negotiating Digital Identities with AI Companions: Motivations, Strategies, and Emotional Outcomes
by: Ma, Renkai, et al.
Published: (2026)
by: Ma, Renkai, et al.
Published: (2026)
User Privacy Harms and Risks in Conversational AI: A Proposed Framework
by: Gumusel, Ece, et al.
Published: (2024)
by: Gumusel, Ece, et al.
Published: (2024)
Tracing Users' Privacy Concerns Across the Lifecycle of a Romantic AI Companion
by: Azam, Kazi Ababil, et al.
Published: (2026)
by: Azam, Kazi Ababil, et al.
Published: (2026)
Embodied Supervision: Haptic Display of Automation Command to Improve Supervisory Performance
by: Gilbert, Alia, et al.
Published: (2024)
by: Gilbert, Alia, et al.
Published: (2024)
Chatting with Confidants or Corporations? Privacy Management with AI Companions
by: Chiu, Hsuen-Chi, et al.
Published: (2026)
by: Chiu, Hsuen-Chi, et al.
Published: (2026)
Designing KRIYA: An AI Companion for Wellbeing Self-Reflection
by: Zhu, Shanshan, et al.
Published: (2026)
by: Zhu, Shanshan, et al.
Published: (2026)
Towards an AI Buddy for every University Student? Exploring Students' Experiences, Attitudes and Motivations towards AI and AI-based Study Companions
by: Moreno, Judit Martinez, et al.
Published: (2026)
by: Moreno, Judit Martinez, et al.
Published: (2026)
SHIELD: LLM-Driven Schema Induction for Predictive Analytics in EV Battery Supply Chain Disruptions
by: Cheng, Zhi-Qi, et al.
Published: (2024)
by: Cheng, Zhi-Qi, et al.
Published: (2024)
Unilateral Relationship Revision Power in Human-AI Companion Interaction
by: Lange, Benjamin
Published: (2026)
by: Lange, Benjamin
Published: (2026)
F.A.C.U.L.: Language-Based Interaction with AI Companions in Gaming
by: Wei, Wenya, et al.
Published: (2025)
by: Wei, Wenya, et al.
Published: (2025)
What Comes After Harm? Mapping Reparative Actions in AI through Justice Frameworks
by: Xiao, Sijia, et al.
Published: (2025)
by: Xiao, Sijia, et al.
Published: (2025)
Beyond the Benefits: A Systematic Review of the Harms and Consequences of Generative AI in Computing Education
by: Bernstein, Seth, et al.
Published: (2025)
by: Bernstein, Seth, et al.
Published: (2025)
AI Generated Child Sexual Abuse Material -- What's the Harm?
by: Ciardha, Caoilte Ó, et al.
Published: (2025)
by: Ciardha, Caoilte Ó, et al.
Published: (2025)
Developers' Experience with Generative AI -- First Insights from an Empirical Mixed-Methods Field Study
by: Brandebusemeyer, Charlotte, et al.
Published: (2025)
by: Brandebusemeyer, Charlotte, et al.
Published: (2025)
"Pragmatic Tools or Empowering Friends?" Discovering and Co-Designing Personality-Aligned AI Writing Companions
by: Wu, Mengke, et al.
Published: (2025)
by: Wu, Mengke, et al.
Published: (2025)
AI and Suicide Prevention: A Cross-Sector Primer
by: Saltz, Emily, et al.
Published: (2026)
by: Saltz, Emily, et al.
Published: (2026)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
Combating Harms of Generative AI in CS1 with Code Review Interviews and a Flipped Classroom
by: Fowles, Peter, et al.
Published: (2026)
by: Fowles, Peter, et al.
Published: (2026)
Exploring Student Behaviors and Motivations using AI TAs with Optional Guardrails
by: Kapoor, Amanpreet, et al.
Published: (2025)
by: Kapoor, Amanpreet, et al.
Published: (2025)
Similar Items
-
Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making
by: Ben-Zion, Ziv, et al.
Published: (2025) -
Harmful Traits of AI Companions
by: Knox, W. Bradley, et al.
Published: (2025) -
Intimacy as Service, Harm as Externality: Critical Perspectives on AI Companion Platform Accountability
by: Eom, Dayeon, et al.
Published: (2026) -
Measuring Machine Companionship: Scale Development and Validation for AI Companions
by: Banks, Jaime
Published: (2025) -
Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions
by: Fan, Xianzhe, et al.
Published: (2024)