AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Naik, Akshat, Quinn, Patrick, Bosch, Guillermo, Gouné, Emma, Zabala, Francisco Javier Campos, Brown, Jason Ross, Young, Edward James |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
por: Wang, Xinyue, et al.
Publicado: (2026)
por: Wang, Xinyue, et al.
Publicado: (2026)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
por: Tang, Wenjie, et al.
Publicado: (2026)
por: Tang, Wenjie, et al.
Publicado: (2026)
Changing the Rules of the Game: Reasoning about Dynamic Phenomena in Multi-Agent Systems
por: Galimullin, Rustam, et al.
Publicado: (2025)
por: Galimullin, Rustam, et al.
Publicado: (2025)
Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines
por: Ray, Aninda
Publicado: (2026)
por: Ray, Aninda
Publicado: (2026)
On The Role of Intentionality in Knowledge Representation: Analyzing Scene Context for Cognitive Agents with a Tiny Language Model
por: Burgess, Mark
Publicado: (2025)
por: Burgess, Mark
Publicado: (2025)
AgentLeak: A Full-Stack Benchmark for Privacy Leakage in Multi-Agent LLM Systems
por: Yagoubi, Faouzi El, et al.
Publicado: (2026)
por: Yagoubi, Faouzi El, et al.
Publicado: (2026)
Umwelt Engineering: Designing the Cognitive Worlds of Linguistic Agents
por: Jehu-Appiah, Rodney
Publicado: (2026)
por: Jehu-Appiah, Rodney
Publicado: (2026)
From Safety Risk to Design Principle: Peer-Preservation in Multi-Agent LLM Systems and Its Implications for Orchestrated Democratic Discourse Analysis
por: Dietrich, Juergen
Publicado: (2026)
por: Dietrich, Juergen
Publicado: (2026)
$γ(3,4)$ `Attention' in Cognitive Agents: Ontology-Free Knowledge Representations With Promise Theoretic Semantics
por: Burgess, Mark
Publicado: (2025)
por: Burgess, Mark
Publicado: (2025)
Privacy as Commodity: MFG-RegretNet for Large-Scale Privacy Trading in Federated Learning
por: Sun, Kangkang, et al.
Publicado: (2026)
por: Sun, Kangkang, et al.
Publicado: (2026)
Policy Cards: Machine-Readable Runtime Governance for Autonomous AI Agents
por: Mavračić, Juraj
Publicado: (2025)
por: Mavračić, Juraj
Publicado: (2025)
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
por: Wu, Shuai, et al.
Publicado: (2026)
por: Wu, Shuai, et al.
Publicado: (2026)
Exploring Design of Multi-Agent LLM Dialogues for Research Ideation
por: Ueda, Keisuke, et al.
Publicado: (2025)
por: Ueda, Keisuke, et al.
Publicado: (2025)
Comparing State-Representations for DEL Model Checking
por: Behnke, Gregor, et al.
Publicado: (2025)
por: Behnke, Gregor, et al.
Publicado: (2025)
FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory
por: Gu, Yingjie, et al.
Publicado: (2026)
por: Gu, Yingjie, et al.
Publicado: (2026)
AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation
por: Costa, Igor
Publicado: (2026)
por: Costa, Igor
Publicado: (2026)
A Super-Learner with Large Language Models for Medical Emergency Advising
por: Aityan, Sergey K., et al.
Publicado: (2025)
por: Aityan, Sergey K., et al.
Publicado: (2025)
Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems: A Neurosymbolic Architecture for Domain-Grounded AI Agents
por: Tuan, Thanh Luong, et al.
Publicado: (2026)
por: Tuan, Thanh Luong, et al.
Publicado: (2026)
N-Agent Ad Hoc Teamwork
por: Wang, Caroline, et al.
Publicado: (2024)
por: Wang, Caroline, et al.
Publicado: (2024)
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
por: Li, Bowen, et al.
Publicado: (2026)
por: Li, Bowen, et al.
Publicado: (2026)
SoccerRef-Agents: Multi-Agent System for Automated Soccer Refereeing
por: Meng, Zi, et al.
Publicado: (2026)
por: Meng, Zi, et al.
Publicado: (2026)
MARS: Multi-Agent Robotic System with Multimodal Large Language Models for Assistive Intelligence
por: Gao, Renjun
Publicado: (2025)
por: Gao, Renjun
Publicado: (2025)
Generating Causal Explanations of Vehicular Agent Behavioural Interactions with Learnt Reward Profiles
por: Howard, Rhys, et al.
Publicado: (2025)
por: Howard, Rhys, et al.
Publicado: (2025)
SDOF: Taming the Alignment Tax in Multi-Agent Orchestration with State-Constrained Dispatch
por: Wang, Zhantao
Publicado: (2026)
por: Wang, Zhantao
Publicado: (2026)
Advancing Multimodal Agent Reasoning with Long-Term Neuro-Symbolic Memory
por: Jiang, Rongjie, et al.
Publicado: (2026)
por: Jiang, Rongjie, et al.
Publicado: (2026)
AgentAtlas: Beyond Outcome Leaderboards for LLM Agents
por: Mazaheri, Parsa, et al.
Publicado: (2026)
por: Mazaheri, Parsa, et al.
Publicado: (2026)
Instruction-Level Weight Shaping: A Framework for Self-Improving AI Agents
por: Costa, Rimom
Publicado: (2025)
por: Costa, Rimom
Publicado: (2025)
Sovereign-OS: A Charter-Governed Operating System for Autonomous AI Agents with Verifiable Fiscal Discipline
por: Yuan, Aojie, et al.
Publicado: (2026)
por: Yuan, Aojie, et al.
Publicado: (2026)
Privacy Preserving Multi Agent Path Finding
por: Lehman, Rotem Lev, et al.
Publicado: (2026)
por: Lehman, Rotem Lev, et al.
Publicado: (2026)
One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents
por: Hong, Yoosung
Publicado: (2026)
por: Hong, Yoosung
Publicado: (2026)
Exploring Robust Multi-Agent Workflows for Environmental Data Management
por: Guan, Boyuan, et al.
Publicado: (2026)
por: Guan, Boyuan, et al.
Publicado: (2026)
Multi-Agent Design Assistant for the Simulation of Inertial Fusion Energy
por: Shachar, Meir H., et al.
Publicado: (2025)
por: Shachar, Meir H., et al.
Publicado: (2025)
Applying Cognitive Design Patterns to General LLM Agents
por: Wray, Robert E., et al.
Publicado: (2025)
por: Wray, Robert E., et al.
Publicado: (2025)
Agent Semantics, Semantic Spacetime, and Graphical Reasoning
por: Burgess, Mark
Publicado: (2025)
por: Burgess, Mark
Publicado: (2025)
Robust and Diverse Multi-Agent Learning via Rational Policy Gradient
por: Lauffer, Niklas, et al.
Publicado: (2025)
por: Lauffer, Niklas, et al.
Publicado: (2025)
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework
por: Wang, Xiaohua, et al.
Publicado: (2026)
por: Wang, Xiaohua, et al.
Publicado: (2026)
Analysing Factorizations of Action-Value Networks for Cooperative Multi-Agent Reinforcement Learning
por: Castellini, Jacopo, et al.
Publicado: (2019)
por: Castellini, Jacopo, et al.
Publicado: (2019)
PRIMA: Operational Patterns for Resilient Multi-Agent Research with Verifiable Identity and Convergent Feedback
por: Annapureddy, Sasank
Publicado: (2026)
por: Annapureddy, Sasank
Publicado: (2026)
ME-IGM: Individual-Global-Max in Maximum Entropy Multi-Agent Reinforcement Learning
por: Chen, Wen-Tse, et al.
Publicado: (2024)
por: Chen, Wen-Tse, et al.
Publicado: (2024)
Evolved Developmental Artificial Neural Networks for Multitasking with Advanced Activity Dependence
por: Zhang, Yintong, et al.
Publicado: (2024)
por: Zhang, Yintong, et al.
Publicado: (2024)
Ejemplares similares
-
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
por: Wang, Xinyue, et al.
Publicado: (2026) -
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
por: Tang, Wenjie, et al.
Publicado: (2026) -
Changing the Rules of the Game: Reasoning about Dynamic Phenomena in Multi-Agent Systems
por: Galimullin, Rustam, et al.
Publicado: (2025) -
Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines
por: Ray, Aninda
Publicado: (2026) -
On The Role of Intentionality in Knowledge Representation: Analyzing Scene Context for Cognitive Agents with a Tiny Language Model
por: Burgess, Mark
Publicado: (2025)