Saved in:
| Main Author: | Milsom, Rory |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.02230 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025)
by: Liu, Zhou, et al.
Published: (2025)
When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs
by: Zeng, Yifan, et al.
Published: (2026)
by: Zeng, Yifan, et al.
Published: (2026)
Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows
by: Chen, Zixin, et al.
Published: (2026)
by: Chen, Zixin, et al.
Published: (2026)
FlowAgent: Achieving Compliance and Flexibility for Workflow Agents
by: Shi, Yuchen, et al.
Published: (2025)
by: Shi, Yuchen, et al.
Published: (2025)
WorkflowPerturb: Calibrated Stress Tests for Evaluating Multi-Agent Workflow Metrics
by: Kanda, Madhav, et al.
Published: (2026)
by: Kanda, Madhav, et al.
Published: (2026)
From Text to Simulation: A Multi-Agent LLM Workflow for Automated Chemical Process Design
by: Tian, Xufei, et al.
Published: (2026)
by: Tian, Xufei, et al.
Published: (2026)
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
by: Li, Chenxin, et al.
Published: (2026)
by: Li, Chenxin, et al.
Published: (2026)
WirelessAgent++: Automated Agentic Workflow Design and Benchmarking for Wireless Networks
by: Tong, Jingwen, et al.
Published: (2026)
by: Tong, Jingwen, et al.
Published: (2026)
Exploring Robust Multi-Agent Workflows for Environmental Data Management
by: Guan, Boyuan, et al.
Published: (2026)
by: Guan, Boyuan, et al.
Published: (2026)
A Declarative Language for Building And Orchestrating LLM-Powered Agent Workflows
by: Daunis, Ivan
Published: (2025)
by: Daunis, Ivan
Published: (2025)
AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks
by: Jiang, Tanqiu, et al.
Published: (2026)
by: Jiang, Tanqiu, et al.
Published: (2026)
Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection
by: Debi, Tanusree, et al.
Published: (2026)
by: Debi, Tanusree, et al.
Published: (2026)
From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
Procedural Knowledge Improves Agentic LLM Workflows
by: Hsiao, Vincent, et al.
Published: (2025)
by: Hsiao, Vincent, et al.
Published: (2025)
Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Multi-Agent Workflows
by: Gharzeddine, Luay, et al.
Published: (2026)
by: Gharzeddine, Luay, et al.
Published: (2026)
Causal Past Logic for Runtime Verification of Distributed LLM Agent Workflows
by: Bollig, Benedikt
Published: (2026)
by: Bollig, Benedikt
Published: (2026)
RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
by: Atinafu, Yonas, et al.
Published: (2026)
by: Atinafu, Yonas, et al.
Published: (2026)
Toward Real-World Table Agents: Capabilities, Workflows, and Design Principles for LLM-based Table Intelligence
by: Tian, Jiaming, et al.
Published: (2025)
by: Tian, Jiaming, et al.
Published: (2025)
Workflows vs Agents for Code Translation
by: Gray, Henry, et al.
Published: (2025)
by: Gray, Henry, et al.
Published: (2025)
From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
by: Yue, Ling, et al.
Published: (2026)
by: Yue, Ling, et al.
Published: (2026)
GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
by: Wang, Jize, et al.
Published: (2026)
by: Wang, Jize, et al.
Published: (2026)
Evaluation and Benchmarking of LLM Agents: A Survey
by: Mohammadi, Mahmoud, et al.
Published: (2025)
by: Mohammadi, Mahmoud, et al.
Published: (2025)
LogicGaze: Benchmarking Causal Consistency in Visual Narratives via Counterfactual Verification
by: Driscoll, Rory, et al.
Published: (2026)
by: Driscoll, Rory, et al.
Published: (2026)
Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
by: Froger, Romain, et al.
Published: (2026)
by: Froger, Romain, et al.
Published: (2026)
Game-theoretic LLM: Agent Workflow for Negotiation Games
by: Hua, Wenyue, et al.
Published: (2024)
by: Hua, Wenyue, et al.
Published: (2024)
AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows
by: Shen, Shuaike, et al.
Published: (2026)
by: Shen, Shuaike, et al.
Published: (2026)
AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks
by: Jorf, Baraa Al, et al.
Published: (2026)
by: Jorf, Baraa Al, et al.
Published: (2026)
WorkflowLLM: Enhancing Workflow Orchestration Capability of Large Language Models
by: Fan, Shengda, et al.
Published: (2024)
by: Fan, Shengda, et al.
Published: (2024)
Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessment
by: Pei, Jiahuan, et al.
Published: (2025)
by: Pei, Jiahuan, et al.
Published: (2025)
EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow
by: Pan, Xiaoyu, et al.
Published: (2025)
by: Pan, Xiaoyu, et al.
Published: (2025)
Automated Multi-Agent Workflows for RTL Design
by: Bhattaram, Amulya, et al.
Published: (2025)
by: Bhattaram, Amulya, et al.
Published: (2025)
Robotouille: An Asynchronous Planning Benchmark for LLM Agents
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
EvoAgentX: An Automated Framework for Evolving Agentic Workflows
by: Wang, Yingxu, et al.
Published: (2025)
by: Wang, Yingxu, et al.
Published: (2025)
Adaptive Multi-Agent Reasoning via Automated Workflow Generation
by: Sami, Humza, et al.
Published: (2025)
by: Sami, Humza, et al.
Published: (2025)
NetGent: Agent-Based Automation of Network Application Workflows
by: Daneshamooz, Jaber, et al.
Published: (2025)
by: Daneshamooz, Jaber, et al.
Published: (2025)
From Intent to Execution: Composing Agentic Workflows with Agent Recommendation
by: Athrey, Kishan, et al.
Published: (2026)
by: Athrey, Kishan, et al.
Published: (2026)
Constrained Process Maps for Multi-Agent Generative AI Workflows
by: Joshi, Ananya, et al.
Published: (2026)
by: Joshi, Ananya, et al.
Published: (2026)
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
by: Wang, Ruipeng, et al.
Published: (2026)
by: Wang, Ruipeng, et al.
Published: (2026)
NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents
by: Chen, Ruolin, et al.
Published: (2025)
by: Chen, Ruolin, et al.
Published: (2025)
Similar Items
-
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025) -
When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs
by: Zeng, Yifan, et al.
Published: (2026) -
Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows
by: Chen, Zixin, et al.
Published: (2026) -
FlowAgent: Achieving Compliance and Flexibility for Workflow Agents
by: Shi, Yuchen, et al.
Published: (2025) -
WorkflowPerturb: Calibrated Stress Tests for Evaluating Multi-Agent Workflow Metrics
by: Kanda, Madhav, et al.
Published: (2026)