Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Gao, Shanshan, Zhou, Liyi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents
di: Sheng, Strick, et al.
Pubblicazione: (2026)
di: Sheng, Strick, et al.
Pubblicazione: (2026)
AI Agent Smart Contract Exploit Generation
di: Gervais, Arthur, et al.
Pubblicazione: (2025)
di: Gervais, Arthur, et al.
Pubblicazione: (2025)
TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
di: Sui, Xingyu, et al.
Pubblicazione: (2026)
di: Sui, Xingyu, et al.
Pubblicazione: (2026)
MoPHES:Leveraging on-device LLMs as Agent for Mobile Psychological Health Evaluation and Support
di: Wei, Xun, et al.
Pubblicazione: (2025)
di: Wei, Xun, et al.
Pubblicazione: (2025)
LATTICE: Evaluating Decision Support Utility of Crypto Agents
di: Chan, Aaron, et al.
Pubblicazione: (2026)
di: Chan, Aaron, et al.
Pubblicazione: (2026)
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
di: Feng, Yunhao, et al.
Pubblicazione: (2026)
di: Feng, Yunhao, et al.
Pubblicazione: (2026)
ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support
di: Chen, Tiantian, et al.
Pubblicazione: (2026)
di: Chen, Tiantian, et al.
Pubblicazione: (2026)
Don't Start What You Can't Finish: A Counterfactual Audit of Support-State Triage in LLM Agents
di: Unlu, Eren
Pubblicazione: (2026)
di: Unlu, Eren
Pubblicazione: (2026)
DeepPsy-Agent: A Stage-Aware and Deep-Thinking Emotional Support Agent System
di: Chen, Kai, et al.
Pubblicazione: (2025)
di: Chen, Kai, et al.
Pubblicazione: (2025)
ChoiceMates: Supporting Unfamiliar Online Decision-Making with Multi-Agent Conversational Interactions
di: Park, Jeongeon, et al.
Pubblicazione: (2023)
di: Park, Jeongeon, et al.
Pubblicazione: (2023)
2-Step Agent: A Framework for the Interaction of a Decision Maker with AI Decision Support
di: Nyberg, Otto, et al.
Pubblicazione: (2026)
di: Nyberg, Otto, et al.
Pubblicazione: (2026)
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
di: Dong, Haonan, et al.
Pubblicazione: (2026)
di: Dong, Haonan, et al.
Pubblicazione: (2026)
SAGE: A Service Agent Graph-guided Evaluation Benchmark
di: Shi, Ling, et al.
Pubblicazione: (2026)
di: Shi, Ling, et al.
Pubblicazione: (2026)
Evaluating Collaborative and Autonomous Agents in Data-Stream-Supported Coordination of Mobile Crowdsourcing
di: Bruns, Ralf, et al.
Pubblicazione: (2024)
di: Bruns, Ralf, et al.
Pubblicazione: (2024)
ANX: Protocol-First Design for AI Agent Interaction with a Supporting 3EX Decoupled Architecture
di: Mingze, Xu
Pubblicazione: (2026)
di: Mingze, Xu
Pubblicazione: (2026)
RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
di: Atinafu, Yonas, et al.
Pubblicazione: (2026)
di: Atinafu, Yonas, et al.
Pubblicazione: (2026)
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
di: Wang, Wanying, et al.
Pubblicazione: (2024)
di: Wang, Wanying, et al.
Pubblicazione: (2024)
Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents
di: Li, Jinyang, et al.
Pubblicazione: (2024)
di: Li, Jinyang, et al.
Pubblicazione: (2024)
OnlineMate: An LLM-Based Multi-Agent Companion System for Cognitive Support in Online Learning
di: Gao, Xian, et al.
Pubblicazione: (2025)
di: Gao, Xian, et al.
Pubblicazione: (2025)
IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents
di: Chen, Rongqian, et al.
Pubblicazione: (2026)
di: Chen, Rongqian, et al.
Pubblicazione: (2026)
Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical Databases
di: Li, Dubai, et al.
Pubblicazione: (2026)
di: Li, Dubai, et al.
Pubblicazione: (2026)
Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning
di: Wu, Yuyang, et al.
Pubblicazione: (2026)
di: Wu, Yuyang, et al.
Pubblicazione: (2026)
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
di: Priyanshu, Aman, et al.
Pubblicazione: (2026)
di: Priyanshu, Aman, et al.
Pubblicazione: (2026)
Do Agents Know What They Can't Do? Evaluating Feasibility Awareness in Tool-Using Agents
di: Cheng, Liang, et al.
Pubblicazione: (2026)
di: Cheng, Liang, et al.
Pubblicazione: (2026)
Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First
di: Liu, Shu, et al.
Pubblicazione: (2025)
di: Liu, Shu, et al.
Pubblicazione: (2025)
SPA-Bench: A Comprehensive Benchmark for SmartPhone Agent Evaluation
di: Chen, Jingxuan, et al.
Pubblicazione: (2024)
di: Chen, Jingxuan, et al.
Pubblicazione: (2024)
macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
di: Yang, Pei, et al.
Pubblicazione: (2025)
di: Yang, Pei, et al.
Pubblicazione: (2025)
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
di: Levy, Ido, et al.
Pubblicazione: (2024)
di: Levy, Ido, et al.
Pubblicazione: (2024)
ArguAgent: AI-Supported Real-Time Grouping for Productive Argumentation in STEM Classrooms
di: Kleiman, Jennifer, et al.
Pubblicazione: (2026)
di: Kleiman, Jennifer, et al.
Pubblicazione: (2026)
Can AI Master Econometrics? Evidence from Econometrics AI Agent on Expert-Level Tasks
di: Chen, Qiang, et al.
Pubblicazione: (2025)
di: Chen, Qiang, et al.
Pubblicazione: (2025)
Toward User Comprehension Supports for LLM Agent Skill Specifications
di: Wen, Zikai Alex
Pubblicazione: (2026)
di: Wen, Zikai Alex
Pubblicazione: (2026)
Human-in-the-Loop Multi-Agent Ventilator Decision Support with Contextual Bandit Preference Learning
di: Li, Sijia, et al.
Pubblicazione: (2026)
di: Li, Sijia, et al.
Pubblicazione: (2026)
Agent-in-the-Loop: A Data Flywheel for Continuous Improvement in LLM-based Customer Support
di: Zhao, Cen Mia, et al.
Pubblicazione: (2025)
di: Zhao, Cen Mia, et al.
Pubblicazione: (2025)
MemoCoder: Automated Function Synthesis using LLM-Supported Agents
di: Jia, Yiping, et al.
Pubblicazione: (2025)
di: Jia, Yiping, et al.
Pubblicazione: (2025)
Evaluating Cognitive Age Alignment in Interactive AI Agents
di: Shen, Yifan, et al.
Pubblicazione: (2026)
di: Shen, Yifan, et al.
Pubblicazione: (2026)
EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents
di: Liu, Yunqi, et al.
Pubblicazione: (2026)
di: Liu, Yunqi, et al.
Pubblicazione: (2026)
ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents
di: Lai, Yuxiang, et al.
Pubblicazione: (2026)
di: Lai, Yuxiang, et al.
Pubblicazione: (2026)
A Novel Task-Driven Method with Evolvable Interactive Agents Using Event Trees for Enhanced Emergency Decision Support
di: Xiao, Xingyu, et al.
Pubblicazione: (2024)
di: Xiao, Xingyu, et al.
Pubblicazione: (2024)
Can Agents Fix Agent Issues?
di: Rahardja, Alfin Wijaya, et al.
Pubblicazione: (2025)
di: Rahardja, Alfin Wijaya, et al.
Pubblicazione: (2025)
Evaluation and Benchmarking of LLM Agents: A Survey
di: Mohammadi, Mahmoud, et al.
Pubblicazione: (2025)
di: Mohammadi, Mahmoud, et al.
Pubblicazione: (2025)
Documenti analoghi
-
When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents
di: Sheng, Strick, et al.
Pubblicazione: (2026) -
AI Agent Smart Contract Exploit Generation
di: Gervais, Arthur, et al.
Pubblicazione: (2025) -
TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
di: Sui, Xingyu, et al.
Pubblicazione: (2026) -
MoPHES:Leveraging on-device LLMs as Agent for Mobile Psychological Health Evaluation and Support
di: Wei, Xun, et al.
Pubblicazione: (2025) -
LATTICE: Evaluating Decision Support Utility of Crypto Agents
di: Chan, Aaron, et al.
Pubblicazione: (2026)