Why Agents Compromise Safety Under Pressure
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Hengle, Tang, Ke |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multimodal Safety Evaluation in Generative Agent Social Simulations
di: Vera, Alhim, et al.
Pubblicazione: (2025)
di: Vera, Alhim, et al.
Pubblicazione: (2025)
Social Catalysts, Not Moral Agents: The Illusion of Alignment in LLM Societies
di: Hu, Yueqing, et al.
Pubblicazione: (2026)
di: Hu, Yueqing, et al.
Pubblicazione: (2026)
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
di: Zhu, Kunlun, et al.
Pubblicazione: (2025)
di: Zhu, Kunlun, et al.
Pubblicazione: (2025)
Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
di: Choi, Younwoo, et al.
Pubblicazione: (2025)
di: Choi, Younwoo, et al.
Pubblicazione: (2025)
Mechanism Plausibility in Generative Agent-Based Modeling
di: Zhao, Patrick, et al.
Pubblicazione: (2026)
di: Zhao, Patrick, et al.
Pubblicazione: (2026)
Multimodal Multi-Agent Empowered Legal Judgment Prediction
di: Kang, Zhaolu, et al.
Pubblicazione: (2026)
di: Kang, Zhaolu, et al.
Pubblicazione: (2026)
Embodied LLM Agents Learn to Cooperate in Organized Teams
di: Guo, Xudong, et al.
Pubblicazione: (2024)
di: Guo, Xudong, et al.
Pubblicazione: (2024)
Law in Silico: Simulating Legal Society with LLM-Based Agents
di: Wang, Yiding, et al.
Pubblicazione: (2025)
di: Wang, Yiding, et al.
Pubblicazione: (2025)
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025)
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025)
The Hidden Strength of Disagreement: Unraveling the Consensus-Diversity Tradeoff in Adaptive Multi-Agent Systems
di: Wu, Zengqing, et al.
Pubblicazione: (2025)
di: Wu, Zengqing, et al.
Pubblicazione: (2025)
Preserving Cultural Identity with Context-Aware Translation Through Multi-Agent AI Systems
di: Anik, Mahfuz Ahmed, et al.
Pubblicazione: (2025)
di: Anik, Mahfuz Ahmed, et al.
Pubblicazione: (2025)
SafeTalkCoach: Diversity-Driven Multi-Agent Simulation for Parent-Teen Health Conversations
di: Tabarsi, Benyamin, et al.
Pubblicazione: (2026)
di: Tabarsi, Benyamin, et al.
Pubblicazione: (2026)
Transforming Competition into Collaboration: The Revolutionary Role of Multi-Agent Systems and Language Models in Modern Organizations
di: Cruz, Carlos Jose Xavier
Pubblicazione: (2024)
di: Cruz, Carlos Jose Xavier
Pubblicazione: (2024)
The High Cost of Incivility: Quantifying Interaction Inefficiency via Multi-Agent Monte Carlo Simulations
di: Mangold, Benedikt
Pubblicazione: (2025)
di: Mangold, Benedikt
Pubblicazione: (2025)
LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models
di: Frisch, Ivar, et al.
Pubblicazione: (2024)
di: Frisch, Ivar, et al.
Pubblicazione: (2024)
I Want to Break Free! Persuasion and Anti-Social Behavior of LLMs in Multi-Agent Settings with Social Hierarchy
di: Campedelli, Gian Maria, et al.
Pubblicazione: (2024)
di: Campedelli, Gian Maria, et al.
Pubblicazione: (2024)
Soft-Label Governance for Distributional Safety in Multi-Agent Systems
di: Aiersilan, Aizierjiang, et al.
Pubblicazione: (2026)
di: Aiersilan, Aizierjiang, et al.
Pubblicazione: (2026)
MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data Generation
di: Li, Mingjin, et al.
Pubblicazione: (2025)
di: Li, Mingjin, et al.
Pubblicazione: (2025)
AI Agent for Education: von Neumann Multi-Agent System Framework
di: Jiang, Yuan-Hao, et al.
Pubblicazione: (2024)
di: Jiang, Yuan-Hao, et al.
Pubblicazione: (2024)
Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems
di: Fukui, Hiroki
Pubblicazione: (2026)
di: Fukui, Hiroki
Pubblicazione: (2026)
When Roles Fail: Epistemic Constraints on Advocate Role Fidelity in LLM-Based Political Statement Analysis
di: Dietrich, Juergen
Pubblicazione: (2026)
di: Dietrich, Juergen
Pubblicazione: (2026)
AI Knows When It's Being Watched: Functional Strategic Action and Contextual Register Modulation in Large Language Models
di: Covas, Vinicius, et al.
Pubblicazione: (2026)
di: Covas, Vinicius, et al.
Pubblicazione: (2026)
An Agentic Evaluation Architecture for Historical Bias Detection in Educational Textbooks
di: Stefan, Gabriel, et al.
Pubblicazione: (2026)
di: Stefan, Gabriel, et al.
Pubblicazione: (2026)
Incorporating LLMs for Large-Scale Urban Complex Mobility Simulation
di: Song, Yu-Lun, et al.
Pubblicazione: (2025)
di: Song, Yu-Lun, et al.
Pubblicazione: (2025)
Among Them: A game-based framework for assessing persuasion capabilities of LLMs
di: Idziejczak, Mateusz, et al.
Pubblicazione: (2025)
di: Idziejczak, Mateusz, et al.
Pubblicazione: (2025)
Escalation Risks from Language Models in Military and Diplomatic Decision-Making
di: Rivera, Juan-Pablo, et al.
Pubblicazione: (2024)
di: Rivera, Juan-Pablo, et al.
Pubblicazione: (2024)
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025)
di: Mushtaq, Abdullah, et al.
Pubblicazione: (2025)
CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs
di: Nasarian, Elham, et al.
Pubblicazione: (2026)
di: Nasarian, Elham, et al.
Pubblicazione: (2026)
MAEBE: Multi-Agent Emergent Behavior Framework
di: Erisken, Sinem, et al.
Pubblicazione: (2025)
di: Erisken, Sinem, et al.
Pubblicazione: (2025)
The Automated but Risky Game: Modeling and Benchmarking Agent-to-Agent Negotiations and Transactions in Consumer Markets
di: Zhu, Shenzhe, et al.
Pubblicazione: (2025)
di: Zhu, Shenzhe, et al.
Pubblicazione: (2025)
Psychologically Enhanced AI Agents
di: Besta, Maciej, et al.
Pubblicazione: (2025)
di: Besta, Maciej, et al.
Pubblicazione: (2025)
Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View
di: Zhang, Jintian, et al.
Pubblicazione: (2023)
di: Zhang, Jintian, et al.
Pubblicazione: (2023)
Shall We Team Up: Exploring Spontaneous Cooperation of Competing LLM Agents
di: Wu, Zengqing, et al.
Pubblicazione: (2024)
di: Wu, Zengqing, et al.
Pubblicazione: (2024)
AI Agents Under EU Law
di: Nannini, Luca, et al.
Pubblicazione: (2026)
di: Nannini, Luca, et al.
Pubblicazione: (2026)
Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory
di: Dai, Gordon, et al.
Pubblicazione: (2024)
di: Dai, Gordon, et al.
Pubblicazione: (2024)
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
di: Zhang, Boxuan, et al.
Pubblicazione: (2026)
di: Zhang, Boxuan, et al.
Pubblicazione: (2026)
Can A Society of Generative Agents Simulate Human Behavior and Inform Public Health Policy? A Case Study on Vaccine Hesitancy
di: Hou, Abe Bohan, et al.
Pubblicazione: (2025)
di: Hou, Abe Bohan, et al.
Pubblicazione: (2025)
AgentCity: Constitutional Governance for Autonomous Agent Economies via Separation of Power
di: Ruan, Anbang, et al.
Pubblicazione: (2026)
di: Ruan, Anbang, et al.
Pubblicazione: (2026)
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
di: Cobben, Pepijn, et al.
Pubblicazione: (2026)
di: Cobben, Pepijn, et al.
Pubblicazione: (2026)
Editing Personality for Large Language Models
di: Mao, Shengyu, et al.
Pubblicazione: (2023)
di: Mao, Shengyu, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Multimodal Safety Evaluation in Generative Agent Social Simulations
di: Vera, Alhim, et al.
Pubblicazione: (2025) -
Social Catalysts, Not Moral Agents: The Illusion of Alignment in LLM Societies
di: Hu, Yueqing, et al.
Pubblicazione: (2026) -
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
di: Zhu, Kunlun, et al.
Pubblicazione: (2025) -
Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
di: Choi, Younwoo, et al.
Pubblicazione: (2025) -
Mechanism Plausibility in Generative Agent-Based Modeling
di: Zhao, Patrick, et al.
Pubblicazione: (2026)