When Agents Disagree With Themselves: Measuring Behavioral Consistency in LLM-Based Agents
Fuente:
arXiv
Saved in:
| Main Author: | Mehta, Aman |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Consistency Amplifies: How Behavioral Variance Shapes Agent Accuracy
by: Mehta, Aman
Published: (2026)
by: Mehta, Aman
Published: (2026)
When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
by: Maryanskyy, Artem
Published: (2026)
by: Maryanskyy, Artem
Published: (2026)
EU-Agent-Bench: Measuring Illegal Behavior of LLM Agents Under EU Law
by: Lichkovski, Ilija, et al.
Published: (2025)
by: Lichkovski, Ilija, et al.
Published: (2025)
When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
by: Najera, Aisha, et al.
Published: (2026)
by: Najera, Aisha, et al.
Published: (2026)
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
by: Priyanshu, Aman, et al.
Published: (2026)
by: Priyanshu, Aman, et al.
Published: (2026)
When Documents Disagree: Measuring Institutional Variation in Transplant Guidance with Retrieval-Augmented Language Models
by: Li, Yubo, et al.
Published: (2026)
by: Li, Yubo, et al.
Published: (2026)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
by: Sun, Bian, et al.
Published: (2026)
by: Sun, Bian, et al.
Published: (2026)
Do Role-Playing Agents Practice What They Preach? Belief-Behavior Consistency in LLM-Based Simulations of Human Trust
by: Mannekote, Amogh, et al.
Published: (2025)
by: Mannekote, Amogh, et al.
Published: (2025)
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
by: Wang, Qiang, et al.
Published: (2025)
by: Wang, Qiang, et al.
Published: (2025)
Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
by: Kasprova, Vira, et al.
Published: (2026)
by: Kasprova, Vira, et al.
Published: (2026)
Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors
by: Wiedermann-Möller, Jonas, et al.
Published: (2026)
by: Wiedermann-Möller, Jonas, et al.
Published: (2026)
When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews
by: Kumar, Sandeep, et al.
Published: (2026)
by: Kumar, Sandeep, et al.
Published: (2026)
Formally Specifying the High-Level Behavior of LLM-Based Agents
by: Crouse, Maxwell, et al.
Published: (2023)
by: Crouse, Maxwell, et al.
Published: (2023)
Agentic Trading: When LLM Agents Meet Financial Markets
by: Xia, Yihan, et al.
Published: (2026)
by: Xia, Yihan, et al.
Published: (2026)
Sequential Behavioral Watermarking for LLM Agents
by: An, Hyeseon, et al.
Published: (2026)
by: An, Hyeseon, et al.
Published: (2026)
EvoXplain: When Machine Learning Models Agree on Predictions but Disagree on Why -- Measuring Mechanistic Multiplicity Across Training Runs
by: Bensmail, Chama
Published: (2025)
by: Bensmail, Chama
Published: (2025)
Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents
by: Ngong, Ivoline, et al.
Published: (2025)
by: Ngong, Ivoline, et al.
Published: (2025)
Self-Attribution Bias: When AI Monitors Go Easy on Themselves
by: Khullar, Dipika, et al.
Published: (2026)
by: Khullar, Dipika, et al.
Published: (2026)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
by: Xu, Wenda, et al.
Published: (2025)
by: Xu, Wenda, et al.
Published: (2025)
When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
by: Wang, Cheng, et al.
Published: (2025)
by: Wang, Cheng, et al.
Published: (2025)
When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered Agents
by: Wang, Zongwei, et al.
Published: (2026)
by: Wang, Zongwei, et al.
Published: (2026)
Agent Drift: Quantifying Behavioral Degradation in Multi-Agent LLM Systems Over Extended Interactions
by: Rath, Abhishek
Published: (2026)
by: Rath, Abhishek
Published: (2026)
Behavioral Consistency Validation for LLM Agents: An Analysis of Trading-Style Switching through Stock-Market Simulation
by: Li, Zeping, et al.
Published: (2026)
by: Li, Zeping, et al.
Published: (2026)
On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents
by: Huang, Jen-tse, et al.
Published: (2024)
by: Huang, Jen-tse, et al.
Published: (2024)
LLM-Based Offline Learning for Embodied Agents via Consistency-Guided Reward Ensemble
by: Lee, Yujeong, et al.
Published: (2024)
by: Lee, Yujeong, et al.
Published: (2024)
When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems
by: Huang, Donghao, et al.
Published: (2026)
by: Huang, Donghao, et al.
Published: (2026)
When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents
by: Sheng, Strick, et al.
Published: (2026)
by: Sheng, Strick, et al.
Published: (2026)
When Developer Aid Becomes Security Debt: A Systematic Analysis of Insecure Behaviors in LLM Coding Agents
by: Kozak, Matous, et al.
Published: (2025)
by: Kozak, Matous, et al.
Published: (2025)
When Numbers Start Talking: Implicit Numerical Coordination Among LLM-Based Agents
by: Buscemi, Alessio, et al.
Published: (2026)
by: Buscemi, Alessio, et al.
Published: (2026)
Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?
by: Zhong, Ziqian, et al.
Published: (2026)
by: Zhong, Ziqian, et al.
Published: (2026)
Adaptive and Explainable AI Agents for Anomaly Detection in Critical IoT Infrastructure using LLM-Enhanced Contextual Reasoning
by: Sharma, Raghav, et al.
Published: (2025)
by: Sharma, Raghav, et al.
Published: (2025)
Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
by: Paglieri, Davide, et al.
Published: (2025)
by: Paglieri, Davide, et al.
Published: (2025)
Agentic Unlearning: When LLM Agent Meets Machine Unlearning
by: Wang, Bin, et al.
Published: (2026)
by: Wang, Bin, et al.
Published: (2026)
Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents
by: Pasternak, Gil, et al.
Published: (2025)
by: Pasternak, Gil, et al.
Published: (2025)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning
by: Wu, Zhengxian, et al.
Published: (2026)
by: Wu, Zhengxian, et al.
Published: (2026)
AgentLens: Visual Analysis for Agent Behaviors in LLM-based Autonomous Systems
by: Lu, Jiaying, et al.
Published: (2024)
by: Lu, Jiaying, et al.
Published: (2024)
Towards Trustworthy Multi-Turn LLM Agents via Behavioral Guidance
by: Gürsun, Gonca
Published: (2025)
by: Gürsun, Gonca
Published: (2025)
When LLMs Disagree: Diagnosing Relevance Filtering Bias and Retrieval Divergence in SDG Search
by: Ingram, William A., et al.
Published: (2025)
by: Ingram, William A., et al.
Published: (2025)
LLM-Based Multi-Agent System for Simulating and Analyzing Marketing and Consumer Behavior
by: Chu, Man-Lin, et al.
Published: (2025)
by: Chu, Man-Lin, et al.
Published: (2025)
Similar Items
-
Consistency Amplifies: How Behavioral Variance Shapes Agent Accuracy
by: Mehta, Aman
Published: (2026) -
When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
by: Maryanskyy, Artem
Published: (2026) -
EU-Agent-Bench: Measuring Illegal Behavior of LLM Agents Under EU Law
by: Lichkovski, Ilija, et al.
Published: (2025) -
When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
by: Najera, Aisha, et al.
Published: (2026) -
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
by: Priyanshu, Aman, et al.
Published: (2026)