Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Cao, Hongliu, Driouich, Ilias, Thomas, Eoin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
di: Cao, Hongliu, et al.
Pubblicazione: (2025)
di: Cao, Hongliu, et al.
Pubblicazione: (2025)
Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework
di: Driouich, Ilias, et al.
Pubblicazione: (2025)
di: Driouich, Ilias, et al.
Pubblicazione: (2025)
Local Model Reconstruction Attacks in Federated Learning and their Uses
di: Driouich, Ilias, et al.
Pubblicazione: (2022)
di: Driouich, Ilias, et al.
Pubblicazione: (2022)
When LLMs Imagine People: A Human-Centered Persona Brainstorm Audit for Bias and Fairness in Creative Applications
di: Cao, Hongliu, et al.
Pubblicazione: (2026)
di: Cao, Hongliu, et al.
Pubblicazione: (2026)
Writing Style Matters: An Examination of Bias and Fairness in Information Retrieval Systems
di: Cao, Hongliu
Pubblicazione: (2024)
di: Cao, Hongliu
Pubblicazione: (2024)
Holistic analysis on the sustainability of Federated Learning across AI product lifecycle
di: Cao, Hongliu
Pubblicazione: (2023)
di: Cao, Hongliu
Pubblicazione: (2023)
Semantic Adapter for Universal Text Embeddings: Diagnosing and Mitigating Negation Blindness to Enhance Universality
di: Cao, Hongliu
Pubblicazione: (2025)
di: Cao, Hongliu
Pubblicazione: (2025)
Recent advances in text embedding: A Comprehensive Review of Top-Performing Methods on the MTEB Benchmark
di: Cao, Hongliu
Pubblicazione: (2024)
di: Cao, Hongliu
Pubblicazione: (2024)
Plan Verification for LLM-Based Embodied Task Completion Agents
di: Hariharan, Ananth, et al.
Pubblicazione: (2025)
di: Hariharan, Ananth, et al.
Pubblicazione: (2025)
Beyond Task Success: Measuring Workflow Fidelity in LLM-Based Agentic Payment Systems
di: Huang, Donghao, et al.
Pubblicazione: (2026)
di: Huang, Donghao, et al.
Pubblicazione: (2026)
I Can't Believe It's Corrupt: Evaluating Corruption in Multi-Agent Governance Systems
di: P, Vedanta S, et al.
Pubblicazione: (2026)
di: P, Vedanta S, et al.
Pubblicazione: (2026)
Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems
di: Akshathala, Sreemaee, et al.
Pubblicazione: (2025)
di: Akshathala, Sreemaee, et al.
Pubblicazione: (2025)
In-Context Prompting Obsoletes Agent Orchestration for Procedural Tasks
di: Dennis, Simon, et al.
Pubblicazione: (2026)
di: Dennis, Simon, et al.
Pubblicazione: (2026)
Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement
di: Forouzandeh, Saman, et al.
Pubblicazione: (2025)
di: Forouzandeh, Saman, et al.
Pubblicazione: (2025)
Beyond Task Success: Behavioral and Representational Diagnostics for WAM and VLA
di: Mai, Hung, et al.
Pubblicazione: (2026)
di: Mai, Hung, et al.
Pubblicazione: (2026)
CORE: Full-Path Evaluation of LLM Agents Beyond Final State
di: Michelakis, Panagiotis, et al.
Pubblicazione: (2025)
di: Michelakis, Panagiotis, et al.
Pubblicazione: (2025)
Beyond Task-Oriented and Chitchat Dialogues: Proactive and Transition-Aware Conversational Agents
di: Yoon, Yejin, et al.
Pubblicazione: (2025)
di: Yoon, Yejin, et al.
Pubblicazione: (2025)
WorkstreamBench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
di: Yen, Thomson, et al.
Pubblicazione: (2026)
di: Yen, Thomson, et al.
Pubblicazione: (2026)
Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents
di: Pasternak, Gil, et al.
Pubblicazione: (2025)
di: Pasternak, Gil, et al.
Pubblicazione: (2025)
TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks
di: Weng, Muyan, et al.
Pubblicazione: (2026)
di: Weng, Muyan, et al.
Pubblicazione: (2026)
Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures
di: Park, Junyoung, et al.
Pubblicazione: (2026)
di: Park, Junyoung, et al.
Pubblicazione: (2026)
SSFF: Investigating LLM Predictive Capabilities for Startup Success through a Multi-Agent Framework with Enhanced Explainability and Performance
di: Wang, Xisen, et al.
Pubblicazione: (2024)
di: Wang, Xisen, et al.
Pubblicazione: (2024)
MineNPC-Task: Task Suite for Memory-Aware Minecraft Agents
di: Doss, Tamil Sudaravan Mohan, et al.
Pubblicazione: (2026)
di: Doss, Tamil Sudaravan Mohan, et al.
Pubblicazione: (2026)
MetaCogAgent: A Metacognitive Multi-Agent LLM Framework with Self-Aware Task Delegation
di: Wang, Chenyu, et al.
Pubblicazione: (2026)
di: Wang, Chenyu, et al.
Pubblicazione: (2026)
On the Importance of Task Complexity in Evaluating LLM-Based Multi-Agent Systems
di: Tang, Bohan, et al.
Pubblicazione: (2025)
di: Tang, Bohan, et al.
Pubblicazione: (2025)
A3: Android Agent Arena for Mobile GUI Agents with Essential-State Procedural Evaluation
di: Chai, Yuxiang, et al.
Pubblicazione: (2025)
di: Chai, Yuxiang, et al.
Pubblicazione: (2025)
AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions
di: Sun, Jingwei, et al.
Pubblicazione: (2026)
di: Sun, Jingwei, et al.
Pubblicazione: (2026)
Can LLM Agents Solve Collaborative Tasks? A Study on Urgency-Aware Planning and Coordination
di: Silva, João Vitor de Carvalho, et al.
Pubblicazione: (2025)
di: Silva, João Vitor de Carvalho, et al.
Pubblicazione: (2025)
Beyond Isolated Tasks: A Framework for Evaluating Coding Agents on Sequential Software Evolution
di: Shastry, KN Ajay, et al.
Pubblicazione: (2026)
di: Shastry, KN Ajay, et al.
Pubblicazione: (2026)
Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents
di: Li, Yifei, et al.
Pubblicazione: (2026)
di: Li, Yifei, et al.
Pubblicazione: (2026)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
di: Shi, Zhichao, et al.
Pubblicazione: (2025)
di: Shi, Zhichao, et al.
Pubblicazione: (2025)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
di: Anupam, Sagnik, et al.
Pubblicazione: (2025)
di: Anupam, Sagnik, et al.
Pubblicazione: (2025)
Zero-shot 3D Map Generation with LLM Agents: A Dual-Agent Architecture for Procedural Content Generation
di: Her, Lim Chien, et al.
Pubblicazione: (2025)
di: Her, Lim Chien, et al.
Pubblicazione: (2025)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
di: Collot, Stephane, et al.
Pubblicazione: (2025)
di: Collot, Stephane, et al.
Pubblicazione: (2025)
LLM Agents Beyond Utility: An Open-Ended Perspective
di: Nachkov, Asen, et al.
Pubblicazione: (2025)
di: Nachkov, Asen, et al.
Pubblicazione: (2025)
Beyond Chemical QA: Evaluating LLM's Chemical Reasoning with Modular Chemical Operations
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
di: Liu, Wenrui, et al.
Pubblicazione: (2025)
di: Liu, Wenrui, et al.
Pubblicazione: (2025)
CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty
di: Kirmayr, Johannes, et al.
Pubblicazione: (2026)
di: Kirmayr, Johannes, et al.
Pubblicazione: (2026)
Distribution-Aware Algorithm Design with LLM Agents
di: Koganti, Saharsh, et al.
Pubblicazione: (2026)
di: Koganti, Saharsh, et al.
Pubblicazione: (2026)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
di: Guo, Zhengkang, et al.
Pubblicazione: (2026)
di: Guo, Zhengkang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
di: Cao, Hongliu, et al.
Pubblicazione: (2025) -
Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework
di: Driouich, Ilias, et al.
Pubblicazione: (2025) -
Local Model Reconstruction Attacks in Federated Learning and their Uses
di: Driouich, Ilias, et al.
Pubblicazione: (2022) -
When LLMs Imagine People: A Human-Centered Persona Brainstorm Audit for Bias and Fairness in Creative Applications
di: Cao, Hongliu, et al.
Pubblicazione: (2026) -
Writing Style Matters: An Examination of Bias and Fairness in Information Retrieval Systems
di: Cao, Hongliu
Pubblicazione: (2024)