Saved in:
| Main Authors: | Fa, Dionizije, Culjak, Marko, Pandza, Bruno, Cupic, Mateo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.21800 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sound Agentic Science Requires Adversarial Experiments
by: Fa, Dionizije, et al.
Published: (2026)
by: Fa, Dionizije, et al.
Published: (2026)
BioAgents: Democratizing Bioinformatics Analysis with Multi-Agent Systems
by: Mehandru, Nikita, et al.
Published: (2025)
by: Mehandru, Nikita, et al.
Published: (2025)
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
by: Lupidi, Alisia, et al.
Published: (2026)
by: Lupidi, Alisia, et al.
Published: (2026)
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
by: Bragg, Jonathan, et al.
Published: (2025)
by: Bragg, Jonathan, et al.
Published: (2025)
Deep Research Bench: Evaluating AI Web Research Agents
by: FutureSearch, et al.
Published: (2025)
by: FutureSearch, et al.
Published: (2025)
MarketBench: Evaluating AI Agents as Market Participants
by: Fradkin, Andrey, et al.
Published: (2026)
by: Fradkin, Andrey, et al.
Published: (2026)
AgentBench: Evaluating LLMs as Agents
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
by: Bazinska, Julia, et al.
Published: (2025)
by: Bazinska, Julia, et al.
Published: (2025)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
by: Zheng, Junhao, et al.
Published: (2025)
by: Zheng, Junhao, et al.
Published: (2025)
WebSuite: Systematically Evaluating Why Web Agents Fail
by: Li, Eric, et al.
Published: (2024)
by: Li, Eric, et al.
Published: (2024)
MAESTRO: Multi-Agent Evaluation Suite for Testing, Reliability, and Observability
by: Ma, Tie, et al.
Published: (2026)
by: Ma, Tie, et al.
Published: (2026)
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
by: Dong, Haonan, et al.
Published: (2026)
by: Dong, Haonan, et al.
Published: (2026)
scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis
by: Workman, Kenny, et al.
Published: (2026)
by: Workman, Kenny, et al.
Published: (2026)
ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
by: Jin, Tengjun, et al.
Published: (2025)
by: Jin, Tengjun, et al.
Published: (2025)
BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows
by: Lau, Elaine, et al.
Published: (2026)
by: Lau, Elaine, et al.
Published: (2026)
TPS-Bench: Evaluating AI Agents' Tool Planning \& Scheduling Abilities in Compounding Tasks
by: Xu, Hanwen, et al.
Published: (2025)
by: Xu, Hanwen, et al.
Published: (2025)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
FIRE-Bench: Evaluating Agents on the Rediscovery of Scientific Insights
by: Wang, Zhen, et al.
Published: (2026)
by: Wang, Zhen, et al.
Published: (2026)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
by: Guo, Zhengkang, et al.
Published: (2026)
by: Guo, Zhengkang, et al.
Published: (2026)
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
by: Levy, Ido, et al.
Published: (2024)
by: Levy, Ido, et al.
Published: (2024)
AgentSearchBench: A Benchmark for AI Agent Search in the Wild
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
by: Nandi, Subhrangshu, et al.
Published: (2025)
by: Nandi, Subhrangshu, et al.
Published: (2025)
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
by: Pereira, Kristen, et al.
Published: (2026)
by: Pereira, Kristen, et al.
Published: (2026)
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
by: Chen, Hui, et al.
Published: (2025)
by: Chen, Hui, et al.
Published: (2025)
An Executable Benchmarking Suite for Tool-Using Agents
by: Zhong, Zhiqing, et al.
Published: (2026)
by: Zhong, Zhiqing, et al.
Published: (2026)
$α^3$-SecBench: A Large-Scale Evaluation Suite of Security, Resilience, and Trust for LLM-based UAV Agents over 6G Networks
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
CocoaBench: Evaluating Unified Digital Agents in the Wild
by: CocoaBench Team, et al.
Published: (2026)
by: CocoaBench Team, et al.
Published: (2026)
HumanStudy-Bench: Towards AI Agent Design for Participant Simulation
by: Liu, Xuan, et al.
Published: (2026)
by: Liu, Xuan, et al.
Published: (2026)
MineNPC-Task: Task Suite for Memory-Aware Minecraft Agents
by: Doss, Tamil Sudaravan Mohan, et al.
Published: (2026)
by: Doss, Tamil Sudaravan Mohan, et al.
Published: (2026)
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
by: Liu, Ruoqi, et al.
Published: (2026)
by: Liu, Ruoqi, et al.
Published: (2026)
TeamBench: Evaluating Agent Coordination under Enforced Role Separation
by: Kim, Yubin, et al.
Published: (2026)
by: Kim, Yubin, et al.
Published: (2026)
AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
by: Shi, Wentao, et al.
Published: (2026)
by: Shi, Wentao, et al.
Published: (2026)
SPA-Bench: A Comprehensive Benchmark for SmartPhone Agent Evaluation
by: Chen, Jingxuan, et al.
Published: (2024)
by: Chen, Jingxuan, et al.
Published: (2024)
EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce
by: Min, Rui, et al.
Published: (2025)
by: Min, Rui, et al.
Published: (2025)
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
by: Xu, Zelai, et al.
Published: (2025)
by: Xu, Zelai, et al.
Published: (2025)
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
by: Wu, Yunze, et al.
Published: (2025)
by: Wu, Yunze, et al.
Published: (2025)
BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments
by: Roohani, Yusuf, et al.
Published: (2024)
by: Roohani, Yusuf, et al.
Published: (2024)
PBT-Bench: Benchmarking AI Agents on Property-Based Testing
by: Jing, Lucas, et al.
Published: (2026)
by: Jing, Lucas, et al.
Published: (2026)
Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AI
by: Kozlova, Anna, et al.
Published: (2026)
by: Kozlova, Anna, et al.
Published: (2026)
EmboCoach-Bench: Benchmarking AI Agents on Developing Embodied Robots
by: Lei, Zixing, et al.
Published: (2026)
by: Lei, Zixing, et al.
Published: (2026)
Similar Items
-
Sound Agentic Science Requires Adversarial Experiments
by: Fa, Dionizije, et al.
Published: (2026) -
BioAgents: Democratizing Bioinformatics Analysis with Multi-Agent Systems
by: Mehandru, Nikita, et al.
Published: (2025) -
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
by: Lupidi, Alisia, et al.
Published: (2026) -
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
by: Bragg, Jonathan, et al.
Published: (2025) -
Deep Research Bench: Evaluating AI Web Research Agents
by: FutureSearch, et al.
Published: (2025)