Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Gurram, Bhaskar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing
von: Chen, Han, et al.
Veröffentlicht: (2026)
von: Chen, Han, et al.
Veröffentlicht: (2026)
AgentPulse: A Continuous Multi-Signal Framework for Evaluating AI Agents in Deployment
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
Agent WARPP: Workflow Adherence via Runtime Parallel Personalization
von: Mazzolenis, Maria Emilia, et al.
Veröffentlicht: (2025)
von: Mazzolenis, Maria Emilia, et al.
Veröffentlicht: (2025)
Towards Computer-Using Personal Agents
von: Bonatti, Piero A., et al.
Veröffentlicht: (2025)
von: Bonatti, Piero A., et al.
Veröffentlicht: (2025)
MALTopic: Multi-Agent LLM Topic Modeling Framework
von: Sharma, Yash
Veröffentlicht: (2026)
von: Sharma, Yash
Veröffentlicht: (2026)
Exploring Design of Multi-Agent LLM Dialogues for Research Ideation
von: Ueda, Keisuke, et al.
Veröffentlicht: (2025)
von: Ueda, Keisuke, et al.
Veröffentlicht: (2025)
SEAR: Schema-Based Evaluation and Routing for LLM Gateways
von: Zhang, Zecheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zecheng, et al.
Veröffentlicht: (2026)
RubikSQL: Lifelong Learning Agentic Knowledge Base as an Industrial NL2SQL System
von: Chen, Zui, et al.
Veröffentlicht: (2025)
von: Chen, Zui, et al.
Veröffentlicht: (2025)
AutoBench: Automating LLM Evaluation through Reciprocal Peer Assessment
von: Loi, Dario, et al.
Veröffentlicht: (2025)
von: Loi, Dario, et al.
Veröffentlicht: (2025)
The Power of Stories: Narrative Priming Shapes How LLM Agents Collaborate and Compete
von: Großmann, Gerrit, et al.
Veröffentlicht: (2025)
von: Großmann, Gerrit, et al.
Veröffentlicht: (2025)
ReDel: A Toolkit for LLM-Powered Recursive Multi-Agent Systems
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
Teams of LLM Agents can Exploit Zero-Day Vulnerabilities
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2024)
A Multi-Memory Segment System for Generating High-Quality Long-Term Memory Content in Agents
von: Zhang, Gaoke, et al.
Veröffentlicht: (2025)
von: Zhang, Gaoke, et al.
Veröffentlicht: (2025)
The Wisdom of Agent Crowds: A Human-AI Interaction Innovation Ignition Framework
von: Yang, Senhao, et al.
Veröffentlicht: (2025)
von: Yang, Senhao, et al.
Veröffentlicht: (2025)
MALLM: Multi-Agent Large Language Models Framework
von: Becker, Jonas, et al.
Veröffentlicht: (2025)
von: Becker, Jonas, et al.
Veröffentlicht: (2025)
Bhakti: A Lightweight Vector Database Management System for Endowing Large Language Models with Semantic Search Capabilities and Memory
von: Wu, Zihao
Veröffentlicht: (2025)
von: Wu, Zihao
Veröffentlicht: (2025)
Large Language Models for Simultaneous Named Entity Extraction and Spelling Correction
von: Whittaker, Edward, et al.
Veröffentlicht: (2024)
von: Whittaker, Edward, et al.
Veröffentlicht: (2024)
Evaluating Personality Traits in Large Language Models: Insights from Psychological Questionnaires
von: Bhandari, Pranav, et al.
Veröffentlicht: (2025)
von: Bhandari, Pranav, et al.
Veröffentlicht: (2025)
compar:IA: The French Government's LLM arena to collect French-language human prompts and preference data
von: Termignon, Lucie, et al.
Veröffentlicht: (2026)
von: Termignon, Lucie, et al.
Veröffentlicht: (2026)
ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents
von: Li, Kenan, et al.
Veröffentlicht: (2026)
von: Li, Kenan, et al.
Veröffentlicht: (2026)
Building the Web for Agents: A Declarative Framework for Agent-Web Interaction
von: Schultze, Sven, et al.
Veröffentlicht: (2025)
von: Schultze, Sven, et al.
Veröffentlicht: (2025)
Literature Review Of Multi-Agent Debate For Problem-Solving
von: Tillmann, Arne
Veröffentlicht: (2025)
von: Tillmann, Arne
Veröffentlicht: (2025)
Tool-RoCo: An Agent-as-Tool Self-organization Large Language Model Benchmark in Multi-robot Cooperation
von: Zhang, Ke, et al.
Veröffentlicht: (2025)
von: Zhang, Ke, et al.
Veröffentlicht: (2025)
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
Talk is Cheap, Communication is Hard: Dynamic Grounding Failures and Repair in Multi-Agent Negotiation
von: Yao, Yiheng, et al.
Veröffentlicht: (2026)
von: Yao, Yiheng, et al.
Veröffentlicht: (2026)
Agentic Automation of BT-RADS Scoring: End-to-End Multi-Agent System for Standardized Brain Tumor Follow-up Assessment
von: Jabal, Mohamed Sobhi, et al.
Veröffentlicht: (2026)
von: Jabal, Mohamed Sobhi, et al.
Veröffentlicht: (2026)
Multi-Agent Systems Powered by Large Language Models: Applications in Swarm Intelligence
von: Jimenez-Romero, Cristian, et al.
Veröffentlicht: (2025)
von: Jimenez-Romero, Cristian, et al.
Veröffentlicht: (2025)
Computational Multi-Agents Society Experiments: Social Modeling Framework Based on Generative Agents
von: Zhang, Hanzhong, et al.
Veröffentlicht: (2025)
von: Zhang, Hanzhong, et al.
Veröffentlicht: (2025)
Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference
von: Dong, Wenxin, et al.
Veröffentlicht: (2026)
von: Dong, Wenxin, et al.
Veröffentlicht: (2026)
Diversification as Risk Minimization
von: Takehi, Rikiya, et al.
Veröffentlicht: (2025)
von: Takehi, Rikiya, et al.
Veröffentlicht: (2025)
Self-Emotion Blended Dialogue Generation in Social Simulation Agents
von: Zhang, Qiang, et al.
Veröffentlicht: (2024)
von: Zhang, Qiang, et al.
Veröffentlicht: (2024)
PAVE: A Cognitive Architecture for Legitimate Violation in Generative Agent Societies
von: Yehia, Ahmad, et al.
Veröffentlicht: (2026)
von: Yehia, Ahmad, et al.
Veröffentlicht: (2026)
The Qualitative Laboratory: Theory Prototyping and Hypothesis Generation with Large Language Models
von: Draelants, Hugues
Veröffentlicht: (2025)
von: Draelants, Hugues
Veröffentlicht: (2025)
CRAwDAD: Causal Reasoning Augmentation with Dual-Agent Debate
von: Vamosi, Finn G., et al.
Veröffentlicht: (2025)
von: Vamosi, Finn G., et al.
Veröffentlicht: (2025)
GSAR: Typed Grounding for Hallucination Detection and Recovery in Multi-Agent LLMs
von: Kamelhar, Federico A.
Veröffentlicht: (2026)
von: Kamelhar, Federico A.
Veröffentlicht: (2026)
PestMA: LLM-based Multi-Agent System for Informed Pest Management
von: Shi, Hongrui, et al.
Veröffentlicht: (2025)
von: Shi, Hongrui, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models in a Complex Hidden Role Game
von: Bauer, Niklas
Veröffentlicht: (2026)
von: Bauer, Niklas
Veröffentlicht: (2026)
Anchor-and-Resume Concession Under Dynamic Pricing for LLM-Augmented Freight Negotiation
von: Nguyen, Hoang, et al.
Veröffentlicht: (2026)
von: Nguyen, Hoang, et al.
Veröffentlicht: (2026)
TAGS: A Test-Time Generalist-Specialist Framework with Retrieval-Augmented Reasoning and Verification
von: Wu, Jianghao, et al.
Veröffentlicht: (2025)
von: Wu, Jianghao, et al.
Veröffentlicht: (2025)
Review of Case-Based Reasoning for LLM Agents: Theoretical Foundations, Architectural Components, and Cognitive Integration
von: Hatalis, Kostas, et al.
Veröffentlicht: (2025)
von: Hatalis, Kostas, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing
von: Chen, Han, et al.
Veröffentlicht: (2026) -
AgentPulse: A Continuous Multi-Signal Framework for Evaluating AI Agents in Deployment
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026) -
Agent WARPP: Workflow Adherence via Runtime Parallel Personalization
von: Mazzolenis, Maria Emilia, et al.
Veröffentlicht: (2025) -
Towards Computer-Using Personal Agents
von: Bonatti, Piero A., et al.
Veröffentlicht: (2025) -
MALTopic: Multi-Agent LLM Topic Modeling Framework
von: Sharma, Yash
Veröffentlicht: (2026)