From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Xinyue, Zhang, Yuanhe, Gong, Zhengshuo, Gao, Haoran, Meng, Fanyu, Zhou, Zhenhong, Sun, Li, Liu, Yang, Su, Sen |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
par: Naik, Akshat, et autres
Publié: (2025)
par: Naik, Akshat, et autres
Publié: (2025)
AgentLeak: A Full-Stack Benchmark for Privacy Leakage in Multi-Agent LLM Systems
par: Yagoubi, Faouzi El, et autres
Publié: (2026)
par: Yagoubi, Faouzi El, et autres
Publié: (2026)
Applying Cognitive Design Patterns to General LLM Agents
par: Wray, Robert E., et autres
Publié: (2025)
par: Wray, Robert E., et autres
Publié: (2025)
Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation
par: Yan, Lingyong, et autres
Publié: (2026)
par: Yan, Lingyong, et autres
Publié: (2026)
Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines
par: Ray, Aninda
Publié: (2026)
par: Ray, Aninda
Publié: (2026)
MEMTIER: Tiered Memory Architecture and Retrieval Bottleneck Analysis for Long-Running Autonomous AI Agents
par: Sidik, Bronislav, et autres
Publié: (2026)
par: Sidik, Bronislav, et autres
Publié: (2026)
ScrapMem: A Bio-inspired Framework for On-device Personalized Agent Memory via Optical Forgetting
par: Chang, Jiale, et autres
Publié: (2026)
par: Chang, Jiale, et autres
Publié: (2026)
Do We Always Need Query-Level Workflows? Rethinking Agentic Workflow Generation for Multi-Agent Systems
par: Wang, Zixu, et autres
Publié: (2026)
par: Wang, Zixu, et autres
Publié: (2026)
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
par: Qi, Jinhu, et autres
Publié: (2026)
par: Qi, Jinhu, et autres
Publié: (2026)
Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults
par: Usman, Rana Muhammad
Publié: (2026)
par: Usman, Rana Muhammad
Publié: (2026)
Exploring Design of Multi-Agent LLM Dialogues for Research Ideation
par: Ueda, Keisuke, et autres
Publié: (2025)
par: Ueda, Keisuke, et autres
Publié: (2025)
Instruction-Level Weight Shaping: A Framework for Self-Improving AI Agents
par: Costa, Rimom
Publié: (2025)
par: Costa, Rimom
Publié: (2025)
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
par: Wang, Yuchen, et autres
Publié: (2026)
par: Wang, Yuchen, et autres
Publié: (2026)
PRIMA: Operational Patterns for Resilient Multi-Agent Research with Verifiable Identity and Convergent Feedback
par: Annapureddy, Sasank
Publié: (2026)
par: Annapureddy, Sasank
Publié: (2026)
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework
par: Wang, Xiaohua, et autres
Publié: (2026)
par: Wang, Xiaohua, et autres
Publié: (2026)
Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay
par: Wang, Xiaohua, et autres
Publié: (2026)
par: Wang, Xiaohua, et autres
Publié: (2026)
Post Hoc Extraction of Pareto Fronts for Continuous Control
par: Thakar, Raghav, et autres
Publié: (2026)
par: Thakar, Raghav, et autres
Publié: (2026)
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
par: Li, Bowen, et autres
Publié: (2026)
par: Li, Bowen, et autres
Publié: (2026)
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
par: Wu, Shuai, et autres
Publié: (2026)
par: Wu, Shuai, et autres
Publié: (2026)
CRAwDAD: Causal Reasoning Augmentation with Dual-Agent Debate
par: Vamosi, Finn G., et autres
Publié: (2025)
par: Vamosi, Finn G., et autres
Publié: (2025)
Agent WARPP: Workflow Adherence via Runtime Parallel Personalization
par: Mazzolenis, Maria Emilia, et autres
Publié: (2025)
par: Mazzolenis, Maria Emilia, et autres
Publié: (2025)
CPEMH: An Agentic Framework for Prompt-Driven Behavior Evaluation and Assurance in Foundation-Model Systems for Mental Health Screening
par: Lorenzoni, Giuliano, et autres
Publié: (2026)
par: Lorenzoni, Giuliano, et autres
Publié: (2026)
GSAR: Typed Grounding for Hallucination Detection and Recovery in Multi-Agent LLMs
par: Kamelhar, Federico A.
Publié: (2026)
par: Kamelhar, Federico A.
Publié: (2026)
Latent Cache Flow: Model-to-Model Communication Without Text
par: Rossi, Maximillian, et autres
Publié: (2026)
par: Rossi, Maximillian, et autres
Publié: (2026)
Generative AI Toolkit -- a framework for increasing the quality of LLM-based applications over their whole life cycle
par: Kohl, Jens, et autres
Publié: (2024)
par: Kohl, Jens, et autres
Publié: (2024)
TRIZ Agents: A Multi-Agent LLM Approach for TRIZ-Based Innovation
par: Szczepanik, Kamil, et autres
Publié: (2025)
par: Szczepanik, Kamil, et autres
Publié: (2025)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
par: Penke, Carolin, et autres
Publié: (2025)
par: Penke, Carolin, et autres
Publié: (2025)
Umwelt Engineering: Designing the Cognitive Worlds of Linguistic Agents
par: Jehu-Appiah, Rodney
Publié: (2026)
par: Jehu-Appiah, Rodney
Publié: (2026)
Controlling Long-Horizon Behavior in Language Model Agents with Explicit State Dynamics
par: Subaharan, Sukesh
Publié: (2026)
par: Subaharan, Sukesh
Publié: (2026)
Fuzzy, Symbolic, and Contextual: Enhancing LLM Instruction via Cognitive Scaffolding
par: Figueiredo, Vanessa
Publié: (2025)
par: Figueiredo, Vanessa
Publié: (2025)
Informed AI Regulation: Comparing the Ethical Frameworks of Leading LLM Chatbots Using an Ethics-Based Audit to Assess Moral Reasoning and Normative Values
par: Chun, Jon, et autres
Publié: (2024)
par: Chun, Jon, et autres
Publié: (2024)
Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation
par: Hartmann, David, et autres
Publié: (2026)
par: Hartmann, David, et autres
Publié: (2026)
Collaborative LLM Agents for C4 Software Architecture Design Automation
par: Szczepanik, Kamil, et autres
Publié: (2025)
par: Szczepanik, Kamil, et autres
Publié: (2025)
Tool-RoCo: An Agent-as-Tool Self-organization Large Language Model Benchmark in Multi-robot Cooperation
par: Zhang, Ke, et autres
Publié: (2025)
par: Zhang, Ke, et autres
Publié: (2025)
Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
par: Hossain, Ariyan, et autres
Publié: (2025)
par: Hossain, Ariyan, et autres
Publié: (2025)
Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays
par: Zmanovskii, Nikita
Publié: (2025)
par: Zmanovskii, Nikita
Publié: (2025)
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
par: Ge, Yuxu
Publié: (2026)
par: Ge, Yuxu
Publié: (2026)
ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
par: Khan, Omer Jauhar
Publié: (2025)
par: Khan, Omer Jauhar
Publié: (2025)
Transforming Computer Security and Public Trust Through the Exploration of Fine-Tuning Large Language Models
par: Crumrine, Garrett, et autres
Publié: (2024)
par: Crumrine, Garrett, et autres
Publié: (2024)
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
par: Lucas, Tom, et autres
Publié: (2026)
par: Lucas, Tom, et autres
Publié: (2026)
Documents similaires
-
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
par: Naik, Akshat, et autres
Publié: (2025) -
AgentLeak: A Full-Stack Benchmark for Privacy Leakage in Multi-Agent LLM Systems
par: Yagoubi, Faouzi El, et autres
Publié: (2026) -
Applying Cognitive Design Patterns to General LLM Agents
par: Wray, Robert E., et autres
Publié: (2025) -
Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation
par: Yan, Lingyong, et autres
Publié: (2026) -
Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines
par: Ray, Aninda
Publié: (2026)