Emotional Cost Functions for AI Safety: Teaching Agents to Feel the Weight of Irreversible Consequences
Fuente:
arXiv
Salvato in:
| Autore principale: | Mopgar, Pandurang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum Learning
di: Li, Xinran, et al.
Pubblicazione: (2025)
di: Li, Xinran, et al.
Pubblicazione: (2025)
Feeling Machines: Ethics, Culture, and the Rise of Emotional AI
di: Chavan, Vivek, et al.
Pubblicazione: (2025)
di: Chavan, Vivek, et al.
Pubblicazione: (2025)
Teaching AI to Feel: A Collaborative, Full-Body Exploration of Emotive Communication
di: Tütüncü, Esen K., et al.
Pubblicazione: (2025)
di: Tütüncü, Esen K., et al.
Pubblicazione: (2025)
AI Safety as Control of Irreversibility: A Systems Framework for Decision-Energy and Sovereignty Boundaries
di: Shu, Wesley, et al.
Pubblicazione: (2026)
di: Shu, Wesley, et al.
Pubblicazione: (2026)
Do LLMs "Feel"? Emotion Circuits Discovery and Control
di: Wang, Chenxi, et al.
Pubblicazione: (2025)
di: Wang, Chenxi, et al.
Pubblicazione: (2025)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
Why We Feel: Breaking Boundaries in Emotional Reasoning with Multimodal Large Language Models
di: Lin, Yuxiang, et al.
Pubblicazione: (2025)
di: Lin, Yuxiang, et al.
Pubblicazione: (2025)
Combining Cost-Constrained Runtime Monitors for AI Safety
di: Hua, Tim Tian, et al.
Pubblicazione: (2025)
di: Hua, Tim Tian, et al.
Pubblicazione: (2025)
What You Feel Is Not What They See: On Predicting Self-Reported Emotion from Third-Party Observer Labels
di: El-Tawil, Yara, et al.
Pubblicazione: (2026)
di: El-Tawil, Yara, et al.
Pubblicazione: (2026)
OOD-MMSafe: Advancing MLLM Safety from Harmful Intent to Hidden Consequences
di: Wen, Ming, et al.
Pubblicazione: (2026)
di: Wen, Ming, et al.
Pubblicazione: (2026)
Surfer-H Meets Holo1: Cost-Efficient Web Agent Powered by Open Weights
di: Andreux, Mathieu, et al.
Pubblicazione: (2025)
di: Andreux, Mathieu, et al.
Pubblicazione: (2025)
Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition
di: Xia, Haiying, et al.
Pubblicazione: (2025)
di: Xia, Haiying, et al.
Pubblicazione: (2025)
SafePro: Evaluating the Safety of Professional-Level AI Agents
di: Zhou, Kaiwen, et al.
Pubblicazione: (2026)
di: Zhou, Kaiwen, et al.
Pubblicazione: (2026)
Building Altruistic and Moral AI Agent with Brain-inspired Emotional Empathy Mechanisms
di: Zhao, Feifei, et al.
Pubblicazione: (2024)
di: Zhao, Feifei, et al.
Pubblicazione: (2024)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
di: Yuan, Youliang, et al.
Pubblicazione: (2024)
di: Yuan, Youliang, et al.
Pubblicazione: (2024)
AgentWall: A Runtime Safety Layer for Local AI Agents
di: Aravind, Ashwin
Pubblicazione: (2026)
di: Aravind, Ashwin
Pubblicazione: (2026)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
di: Li, Miles Q., et al.
Pubblicazione: (2026)
di: Li, Miles Q., et al.
Pubblicazione: (2026)
Safeguarding AI Agents: Developing and Analyzing Safety Architectures
di: Domkundwar, Ishaan, et al.
Pubblicazione: (2024)
di: Domkundwar, Ishaan, et al.
Pubblicazione: (2024)
Offline Training of Language Model Agents with Functions as Learnable Weights
di: Zhang, Shaokun, et al.
Pubblicazione: (2024)
di: Zhang, Shaokun, et al.
Pubblicazione: (2024)
LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing
di: Di Palma, Dario, et al.
Pubblicazione: (2025)
di: Di Palma, Dario, et al.
Pubblicazione: (2025)
Improving the Safety and Trustworthiness of Medical AI via Multi-Agent Evaluation Loops
di: Ghafoor, Zainab, et al.
Pubblicazione: (2026)
di: Ghafoor, Zainab, et al.
Pubblicazione: (2026)
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
di: Yang, Chenglin
Pubblicazione: (2026)
di: Yang, Chenglin
Pubblicazione: (2026)
Do You Feel Comfortable? Detecting Hidden Conversational Escalation in AI Chatbots
di: Park, Jihyung, et al.
Pubblicazione: (2025)
di: Park, Jihyung, et al.
Pubblicazione: (2025)
Limits to AI Growth: The Ecological and Social Consequences of Scaling
di: Bhardwaj, Eshta, et al.
Pubblicazione: (2025)
di: Bhardwaj, Eshta, et al.
Pubblicazione: (2025)
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
di: Yin, Lu, et al.
Pubblicazione: (2023)
di: Yin, Lu, et al.
Pubblicazione: (2023)
BeSafe-Bench: Unveiling Behavioral Safety Risks of Situated Agents in Functional Environments
di: Li, Yuxuan, et al.
Pubblicazione: (2026)
di: Li, Yuxuan, et al.
Pubblicazione: (2026)
Robust Agent Compensation (RAC): Teaching AI Agents to Compensate
di: Perera, Srinath, et al.
Pubblicazione: (2026)
di: Perera, Srinath, et al.
Pubblicazione: (2026)
AI with Emotions: Exploring Emotional Expressions in Large Language Models
di: Ishikawa, Shin-nosuke, et al.
Pubblicazione: (2025)
di: Ishikawa, Shin-nosuke, et al.
Pubblicazione: (2025)
Emotional RAG: Enhancing Role-Playing Agents through Emotional Retrieval
di: Huang, Le, et al.
Pubblicazione: (2024)
di: Huang, Le, et al.
Pubblicazione: (2024)
TeachAnything: A Multimodal Crowdsourcing Platform for Training Embodied AI Agents in Symmetrical Reality
di: Liu, Zidong, et al.
Pubblicazione: (2026)
di: Liu, Zidong, et al.
Pubblicazione: (2026)
Simulating Human Behavior with the Psychological-mechanism Agent: Integrating Feeling, Thought, and Action
di: Dong, Qing, et al.
Pubblicazione: (2025)
di: Dong, Qing, et al.
Pubblicazione: (2025)
Solution for Emotion Prediction Competition of Workshop on Emotionally and Culturally Intelligent AI
di: Xu, Shengdong, et al.
Pubblicazione: (2024)
di: Xu, Shengdong, et al.
Pubblicazione: (2024)
DeepKnown-Guard: A Proprietary Model-Based Safety Response Framework for AI Agents
di: Li, Qi, et al.
Pubblicazione: (2025)
di: Li, Qi, et al.
Pubblicazione: (2025)
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
di: Overman, William, et al.
Pubblicazione: (2025)
di: Overman, William, et al.
Pubblicazione: (2025)
Your Robot Will Feel You Now: Empathy in Robots and Embodied Agents
di: Lim, Angelica, et al.
Pubblicazione: (2026)
di: Lim, Angelica, et al.
Pubblicazione: (2026)
Generative AI Agents in Autonomous Machines: A Safety Perspective
di: Jabbour, Jason, et al.
Pubblicazione: (2024)
di: Jabbour, Jason, et al.
Pubblicazione: (2024)
The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
di: Staufer, Leon, et al.
Pubblicazione: (2026)
di: Staufer, Leon, et al.
Pubblicazione: (2026)
Agentic AI Ecosystems in Higher Education: A Perspective on AI Agents to Emerging Inclusive, Agentic Multi-Agent AI Framework for Learning, Teaching and Institutional Intelligence
di: Sudarshan, Vidya K, et al.
Pubblicazione: (2026)
di: Sudarshan, Vidya K, et al.
Pubblicazione: (2026)
The Internet of Physical AI Agents: Interoperability, Longevity, and the Cost of Getting It Wrong
di: Morabito, Roberto, et al.
Pubblicazione: (2026)
di: Morabito, Roberto, et al.
Pubblicazione: (2026)
Thermodynamic Irreversibility of Training Algorithms
di: Ziyin, Liu, et al.
Pubblicazione: (2026)
di: Ziyin, Liu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum Learning
di: Li, Xinran, et al.
Pubblicazione: (2025) -
Feeling Machines: Ethics, Culture, and the Rise of Emotional AI
di: Chavan, Vivek, et al.
Pubblicazione: (2025) -
Teaching AI to Feel: A Collaborative, Full-Body Exploration of Emotive Communication
di: Tütüncü, Esen K., et al.
Pubblicazione: (2025) -
AI Safety as Control of Irreversibility: A Systems Framework for Decision-Energy and Sovereignty Boundaries
di: Shu, Wesley, et al.
Pubblicazione: (2026) -
Do LLMs "Feel"? Emotion Circuits Discovery and Control
di: Wang, Chenxi, et al.
Pubblicazione: (2025)