Emotional Cost Functions for AI Safety: Teaching Agents to Feel the Weight of Irreversible Consequences
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Mopgar, Pandurang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum Learning
von: Li, Xinran, et al.
Veröffentlicht: (2025)
von: Li, Xinran, et al.
Veröffentlicht: (2025)
Feeling Machines: Ethics, Culture, and the Rise of Emotional AI
von: Chavan, Vivek, et al.
Veröffentlicht: (2025)
von: Chavan, Vivek, et al.
Veröffentlicht: (2025)
Teaching AI to Feel: A Collaborative, Full-Body Exploration of Emotive Communication
von: Tütüncü, Esen K., et al.
Veröffentlicht: (2025)
von: Tütüncü, Esen K., et al.
Veröffentlicht: (2025)
AI Safety as Control of Irreversibility: A Systems Framework for Decision-Energy and Sovereignty Boundaries
von: Shu, Wesley, et al.
Veröffentlicht: (2026)
von: Shu, Wesley, et al.
Veröffentlicht: (2026)
Do LLMs "Feel"? Emotion Circuits Discovery and Control
von: Wang, Chenxi, et al.
Veröffentlicht: (2025)
von: Wang, Chenxi, et al.
Veröffentlicht: (2025)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
Why We Feel: Breaking Boundaries in Emotional Reasoning with Multimodal Large Language Models
von: Lin, Yuxiang, et al.
Veröffentlicht: (2025)
von: Lin, Yuxiang, et al.
Veröffentlicht: (2025)
Combining Cost-Constrained Runtime Monitors for AI Safety
von: Hua, Tim Tian, et al.
Veröffentlicht: (2025)
von: Hua, Tim Tian, et al.
Veröffentlicht: (2025)
What You Feel Is Not What They See: On Predicting Self-Reported Emotion from Third-Party Observer Labels
von: El-Tawil, Yara, et al.
Veröffentlicht: (2026)
von: El-Tawil, Yara, et al.
Veröffentlicht: (2026)
OOD-MMSafe: Advancing MLLM Safety from Harmful Intent to Hidden Consequences
von: Wen, Ming, et al.
Veröffentlicht: (2026)
von: Wen, Ming, et al.
Veröffentlicht: (2026)
Surfer-H Meets Holo1: Cost-Efficient Web Agent Powered by Open Weights
von: Andreux, Mathieu, et al.
Veröffentlicht: (2025)
von: Andreux, Mathieu, et al.
Veröffentlicht: (2025)
Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition
von: Xia, Haiying, et al.
Veröffentlicht: (2025)
von: Xia, Haiying, et al.
Veröffentlicht: (2025)
SafePro: Evaluating the Safety of Professional-Level AI Agents
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2026)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2026)
Building Altruistic and Moral AI Agent with Brain-inspired Emotional Empathy Mechanisms
von: Zhao, Feifei, et al.
Veröffentlicht: (2024)
von: Zhao, Feifei, et al.
Veröffentlicht: (2024)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
von: Yuan, Youliang, et al.
Veröffentlicht: (2024)
von: Yuan, Youliang, et al.
Veröffentlicht: (2024)
AgentWall: A Runtime Safety Layer for Local AI Agents
von: Aravind, Ashwin
Veröffentlicht: (2026)
von: Aravind, Ashwin
Veröffentlicht: (2026)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
von: Li, Miles Q., et al.
Veröffentlicht: (2026)
von: Li, Miles Q., et al.
Veröffentlicht: (2026)
Safeguarding AI Agents: Developing and Analyzing Safety Architectures
von: Domkundwar, Ishaan, et al.
Veröffentlicht: (2024)
von: Domkundwar, Ishaan, et al.
Veröffentlicht: (2024)
Offline Training of Language Model Agents with Functions as Learnable Weights
von: Zhang, Shaokun, et al.
Veröffentlicht: (2024)
von: Zhang, Shaokun, et al.
Veröffentlicht: (2024)
LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing
von: Di Palma, Dario, et al.
Veröffentlicht: (2025)
von: Di Palma, Dario, et al.
Veröffentlicht: (2025)
Improving the Safety and Trustworthiness of Medical AI via Multi-Agent Evaluation Loops
von: Ghafoor, Zainab, et al.
Veröffentlicht: (2026)
von: Ghafoor, Zainab, et al.
Veröffentlicht: (2026)
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
von: Yang, Chenglin
Veröffentlicht: (2026)
von: Yang, Chenglin
Veröffentlicht: (2026)
Do You Feel Comfortable? Detecting Hidden Conversational Escalation in AI Chatbots
von: Park, Jihyung, et al.
Veröffentlicht: (2025)
von: Park, Jihyung, et al.
Veröffentlicht: (2025)
Limits to AI Growth: The Ecological and Social Consequences of Scaling
von: Bhardwaj, Eshta, et al.
Veröffentlicht: (2025)
von: Bhardwaj, Eshta, et al.
Veröffentlicht: (2025)
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
von: Yin, Lu, et al.
Veröffentlicht: (2023)
von: Yin, Lu, et al.
Veröffentlicht: (2023)
BeSafe-Bench: Unveiling Behavioral Safety Risks of Situated Agents in Functional Environments
von: Li, Yuxuan, et al.
Veröffentlicht: (2026)
von: Li, Yuxuan, et al.
Veröffentlicht: (2026)
Robust Agent Compensation (RAC): Teaching AI Agents to Compensate
von: Perera, Srinath, et al.
Veröffentlicht: (2026)
von: Perera, Srinath, et al.
Veröffentlicht: (2026)
AI with Emotions: Exploring Emotional Expressions in Large Language Models
von: Ishikawa, Shin-nosuke, et al.
Veröffentlicht: (2025)
von: Ishikawa, Shin-nosuke, et al.
Veröffentlicht: (2025)
Emotional RAG: Enhancing Role-Playing Agents through Emotional Retrieval
von: Huang, Le, et al.
Veröffentlicht: (2024)
von: Huang, Le, et al.
Veröffentlicht: (2024)
TeachAnything: A Multimodal Crowdsourcing Platform for Training Embodied AI Agents in Symmetrical Reality
von: Liu, Zidong, et al.
Veröffentlicht: (2026)
von: Liu, Zidong, et al.
Veröffentlicht: (2026)
Simulating Human Behavior with the Psychological-mechanism Agent: Integrating Feeling, Thought, and Action
von: Dong, Qing, et al.
Veröffentlicht: (2025)
von: Dong, Qing, et al.
Veröffentlicht: (2025)
Solution for Emotion Prediction Competition of Workshop on Emotionally and Culturally Intelligent AI
von: Xu, Shengdong, et al.
Veröffentlicht: (2024)
von: Xu, Shengdong, et al.
Veröffentlicht: (2024)
DeepKnown-Guard: A Proprietary Model-Based Safety Response Framework for AI Agents
von: Li, Qi, et al.
Veröffentlicht: (2025)
von: Li, Qi, et al.
Veröffentlicht: (2025)
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
von: Overman, William, et al.
Veröffentlicht: (2025)
von: Overman, William, et al.
Veröffentlicht: (2025)
Your Robot Will Feel You Now: Empathy in Robots and Embodied Agents
von: Lim, Angelica, et al.
Veröffentlicht: (2026)
von: Lim, Angelica, et al.
Veröffentlicht: (2026)
Generative AI Agents in Autonomous Machines: A Safety Perspective
von: Jabbour, Jason, et al.
Veröffentlicht: (2024)
von: Jabbour, Jason, et al.
Veröffentlicht: (2024)
The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
von: Staufer, Leon, et al.
Veröffentlicht: (2026)
von: Staufer, Leon, et al.
Veröffentlicht: (2026)
Agentic AI Ecosystems in Higher Education: A Perspective on AI Agents to Emerging Inclusive, Agentic Multi-Agent AI Framework for Learning, Teaching and Institutional Intelligence
von: Sudarshan, Vidya K, et al.
Veröffentlicht: (2026)
von: Sudarshan, Vidya K, et al.
Veröffentlicht: (2026)
The Internet of Physical AI Agents: Interoperability, Longevity, and the Cost of Getting It Wrong
von: Morabito, Roberto, et al.
Veröffentlicht: (2026)
von: Morabito, Roberto, et al.
Veröffentlicht: (2026)
Thermodynamic Irreversibility of Training Algorithms
von: Ziyin, Liu, et al.
Veröffentlicht: (2026)
von: Ziyin, Liu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum Learning
von: Li, Xinran, et al.
Veröffentlicht: (2025) -
Feeling Machines: Ethics, Culture, and the Rise of Emotional AI
von: Chavan, Vivek, et al.
Veröffentlicht: (2025) -
Teaching AI to Feel: A Collaborative, Full-Body Exploration of Emotive Communication
von: Tütüncü, Esen K., et al.
Veröffentlicht: (2025) -
AI Safety as Control of Irreversibility: A Systems Framework for Decision-Energy and Sovereignty Boundaries
von: Shu, Wesley, et al.
Veröffentlicht: (2026) -
Do LLMs "Feel"? Emotion Circuits Discovery and Control
von: Wang, Chenxi, et al.
Veröffentlicht: (2025)