Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cao, Hongliu, Driouich, Ilias, Thomas, Eoin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
von: Cao, Hongliu, et al.
Veröffentlicht: (2025)
von: Cao, Hongliu, et al.
Veröffentlicht: (2025)
Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework
von: Driouich, Ilias, et al.
Veröffentlicht: (2025)
von: Driouich, Ilias, et al.
Veröffentlicht: (2025)
Local Model Reconstruction Attacks in Federated Learning and their Uses
von: Driouich, Ilias, et al.
Veröffentlicht: (2022)
von: Driouich, Ilias, et al.
Veröffentlicht: (2022)
When LLMs Imagine People: A Human-Centered Persona Brainstorm Audit for Bias and Fairness in Creative Applications
von: Cao, Hongliu, et al.
Veröffentlicht: (2026)
von: Cao, Hongliu, et al.
Veröffentlicht: (2026)
Writing Style Matters: An Examination of Bias and Fairness in Information Retrieval Systems
von: Cao, Hongliu
Veröffentlicht: (2024)
von: Cao, Hongliu
Veröffentlicht: (2024)
Holistic analysis on the sustainability of Federated Learning across AI product lifecycle
von: Cao, Hongliu
Veröffentlicht: (2023)
von: Cao, Hongliu
Veröffentlicht: (2023)
Semantic Adapter for Universal Text Embeddings: Diagnosing and Mitigating Negation Blindness to Enhance Universality
von: Cao, Hongliu
Veröffentlicht: (2025)
von: Cao, Hongliu
Veröffentlicht: (2025)
Recent advances in text embedding: A Comprehensive Review of Top-Performing Methods on the MTEB Benchmark
von: Cao, Hongliu
Veröffentlicht: (2024)
von: Cao, Hongliu
Veröffentlicht: (2024)
Plan Verification for LLM-Based Embodied Task Completion Agents
von: Hariharan, Ananth, et al.
Veröffentlicht: (2025)
von: Hariharan, Ananth, et al.
Veröffentlicht: (2025)
Beyond Task Success: Measuring Workflow Fidelity in LLM-Based Agentic Payment Systems
von: Huang, Donghao, et al.
Veröffentlicht: (2026)
von: Huang, Donghao, et al.
Veröffentlicht: (2026)
I Can't Believe It's Corrupt: Evaluating Corruption in Multi-Agent Governance Systems
von: P, Vedanta S, et al.
Veröffentlicht: (2026)
von: P, Vedanta S, et al.
Veröffentlicht: (2026)
Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems
von: Akshathala, Sreemaee, et al.
Veröffentlicht: (2025)
von: Akshathala, Sreemaee, et al.
Veröffentlicht: (2025)
In-Context Prompting Obsoletes Agent Orchestration for Procedural Tasks
von: Dennis, Simon, et al.
Veröffentlicht: (2026)
von: Dennis, Simon, et al.
Veröffentlicht: (2026)
Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement
von: Forouzandeh, Saman, et al.
Veröffentlicht: (2025)
von: Forouzandeh, Saman, et al.
Veröffentlicht: (2025)
Beyond Task Success: Behavioral and Representational Diagnostics for WAM and VLA
von: Mai, Hung, et al.
Veröffentlicht: (2026)
von: Mai, Hung, et al.
Veröffentlicht: (2026)
CORE: Full-Path Evaluation of LLM Agents Beyond Final State
von: Michelakis, Panagiotis, et al.
Veröffentlicht: (2025)
von: Michelakis, Panagiotis, et al.
Veröffentlicht: (2025)
Beyond Task-Oriented and Chitchat Dialogues: Proactive and Transition-Aware Conversational Agents
von: Yoon, Yejin, et al.
Veröffentlicht: (2025)
von: Yoon, Yejin, et al.
Veröffentlicht: (2025)
WorkstreamBench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
von: Yen, Thomson, et al.
Veröffentlicht: (2026)
von: Yen, Thomson, et al.
Veröffentlicht: (2026)
Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents
von: Pasternak, Gil, et al.
Veröffentlicht: (2025)
von: Pasternak, Gil, et al.
Veröffentlicht: (2025)
TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks
von: Weng, Muyan, et al.
Veröffentlicht: (2026)
von: Weng, Muyan, et al.
Veröffentlicht: (2026)
Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
SSFF: Investigating LLM Predictive Capabilities for Startup Success through a Multi-Agent Framework with Enhanced Explainability and Performance
von: Wang, Xisen, et al.
Veröffentlicht: (2024)
von: Wang, Xisen, et al.
Veröffentlicht: (2024)
MineNPC-Task: Task Suite for Memory-Aware Minecraft Agents
von: Doss, Tamil Sudaravan Mohan, et al.
Veröffentlicht: (2026)
von: Doss, Tamil Sudaravan Mohan, et al.
Veröffentlicht: (2026)
MetaCogAgent: A Metacognitive Multi-Agent LLM Framework with Self-Aware Task Delegation
von: Wang, Chenyu, et al.
Veröffentlicht: (2026)
von: Wang, Chenyu, et al.
Veröffentlicht: (2026)
On the Importance of Task Complexity in Evaluating LLM-Based Multi-Agent Systems
von: Tang, Bohan, et al.
Veröffentlicht: (2025)
von: Tang, Bohan, et al.
Veröffentlicht: (2025)
A3: Android Agent Arena for Mobile GUI Agents with Essential-State Procedural Evaluation
von: Chai, Yuxiang, et al.
Veröffentlicht: (2025)
von: Chai, Yuxiang, et al.
Veröffentlicht: (2025)
AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions
von: Sun, Jingwei, et al.
Veröffentlicht: (2026)
von: Sun, Jingwei, et al.
Veröffentlicht: (2026)
Can LLM Agents Solve Collaborative Tasks? A Study on Urgency-Aware Planning and Coordination
von: Silva, João Vitor de Carvalho, et al.
Veröffentlicht: (2025)
von: Silva, João Vitor de Carvalho, et al.
Veröffentlicht: (2025)
Beyond Isolated Tasks: A Framework for Evaluating Coding Agents on Sequential Software Evolution
von: Shastry, KN Ajay, et al.
Veröffentlicht: (2026)
von: Shastry, KN Ajay, et al.
Veröffentlicht: (2026)
Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents
von: Li, Yifei, et al.
Veröffentlicht: (2026)
von: Li, Yifei, et al.
Veröffentlicht: (2026)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
von: Shi, Zhichao, et al.
Veröffentlicht: (2025)
von: Shi, Zhichao, et al.
Veröffentlicht: (2025)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
von: Anupam, Sagnik, et al.
Veröffentlicht: (2025)
von: Anupam, Sagnik, et al.
Veröffentlicht: (2025)
Zero-shot 3D Map Generation with LLM Agents: A Dual-Agent Architecture for Procedural Content Generation
von: Her, Lim Chien, et al.
Veröffentlicht: (2025)
von: Her, Lim Chien, et al.
Veröffentlicht: (2025)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
von: Collot, Stephane, et al.
Veröffentlicht: (2025)
von: Collot, Stephane, et al.
Veröffentlicht: (2025)
LLM Agents Beyond Utility: An Open-Ended Perspective
von: Nachkov, Asen, et al.
Veröffentlicht: (2025)
von: Nachkov, Asen, et al.
Veröffentlicht: (2025)
Beyond Chemical QA: Evaluating LLM's Chemical Reasoning with Modular Chemical Operations
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty
von: Kirmayr, Johannes, et al.
Veröffentlicht: (2026)
von: Kirmayr, Johannes, et al.
Veröffentlicht: (2026)
Distribution-Aware Algorithm Design with LLM Agents
von: Koganti, Saharsh, et al.
Veröffentlicht: (2026)
von: Koganti, Saharsh, et al.
Veröffentlicht: (2026)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
von: Guo, Zhengkang, et al.
Veröffentlicht: (2026)
von: Guo, Zhengkang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
von: Cao, Hongliu, et al.
Veröffentlicht: (2025) -
Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework
von: Driouich, Ilias, et al.
Veröffentlicht: (2025) -
Local Model Reconstruction Attacks in Federated Learning and their Uses
von: Driouich, Ilias, et al.
Veröffentlicht: (2022) -
When LLMs Imagine People: A Human-Centered Persona Brainstorm Audit for Bias and Fairness in Creative Applications
von: Cao, Hongliu, et al.
Veröffentlicht: (2026) -
Writing Style Matters: An Examination of Bias and Fairness in Information Retrieval Systems
von: Cao, Hongliu
Veröffentlicht: (2024)