Salvato in:
| Autori principali: | Mohammadi, Mahmoud, Li, Yipeng, Lo, Jane, Yip, Wendy |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2507.21504 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Survey on Evaluation of LLM-based Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
AUTOCT: Automating Interpretable Clinical Trial Prediction with LLM Agents
di: Liu, Fengze, et al.
Pubblicazione: (2025)
di: Liu, Fengze, et al.
Pubblicazione: (2025)
TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks
di: Weng, Muyan, et al.
Pubblicazione: (2026)
di: Weng, Muyan, et al.
Pubblicazione: (2026)
The Evaluation Game: Beyond Static LLM Benchmarking
di: Wang, Paul, et al.
Pubblicazione: (2026)
di: Wang, Paul, et al.
Pubblicazione: (2026)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2025)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2025)
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
di: Thaman, Kunvar
Pubblicazione: (2026)
di: Thaman, Kunvar
Pubblicazione: (2026)
RxEval: A Prescription-Level Benchmark for Evaluating LLM Medication Recommendation
di: Chen, Shuhao, et al.
Pubblicazione: (2026)
di: Chen, Shuhao, et al.
Pubblicazione: (2026)
A Survey on Code Generation with LLM-based Agents
di: Dong, Yihong, et al.
Pubblicazione: (2025)
di: Dong, Yihong, et al.
Pubblicazione: (2025)
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
di: Potham, Ram
Pubblicazione: (2025)
di: Potham, Ram
Pubblicazione: (2025)
DataSciBench: An LLM Agent Benchmark for Data Science
di: Zhang, Dan, et al.
Pubblicazione: (2025)
di: Zhang, Dan, et al.
Pubblicazione: (2025)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
di: Anupam, Sagnik, et al.
Pubblicazione: (2025)
di: Anupam, Sagnik, et al.
Pubblicazione: (2025)
MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness
di: Hathidara, Ashutosh, et al.
Pubblicazione: (2026)
di: Hathidara, Ashutosh, et al.
Pubblicazione: (2026)
Evaluating Long-Context Reasoning in LLM-Based WebAgents
di: Chung, Andy, et al.
Pubblicazione: (2025)
di: Chung, Andy, et al.
Pubblicazione: (2025)
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
di: Yoa, Seungdong, et al.
Pubblicazione: (2026)
di: Yoa, Seungdong, et al.
Pubblicazione: (2026)
Integrating Temporal and Structural Context in Graph Transformers for Relational Deep Learning
di: Lachi, Divyansha, et al.
Pubblicazione: (2025)
di: Lachi, Divyansha, et al.
Pubblicazione: (2025)
BERT-LSH: Reducing Absolute Compute For Attention
di: Li, Zezheng, et al.
Pubblicazione: (2024)
di: Li, Zezheng, et al.
Pubblicazione: (2024)
SmartPlay: A Benchmark for LLMs as Intelligent Agents
di: Wu, Yue, et al.
Pubblicazione: (2023)
di: Wu, Yue, et al.
Pubblicazione: (2023)
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
di: Li, Xiaozhe, et al.
Pubblicazione: (2025)
di: Li, Xiaozhe, et al.
Pubblicazione: (2025)
The LLM Data Auditor: A Metric-oriented Survey on Quality and Trustworthiness in Evaluating Synthetic Data
di: Zhang, Kaituo, et al.
Pubblicazione: (2026)
di: Zhang, Kaituo, et al.
Pubblicazione: (2026)
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
di: Jiang, Yixing, et al.
Pubblicazione: (2025)
di: Jiang, Yixing, et al.
Pubblicazione: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024)
di: Tan, Sijun, et al.
Pubblicazione: (2024)
MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents
di: Rosen, Simon, et al.
Pubblicazione: (2026)
di: Rosen, Simon, et al.
Pubblicazione: (2026)
On the Importance of Task Complexity in Evaluating LLM-Based Multi-Agent Systems
di: Tang, Bohan, et al.
Pubblicazione: (2025)
di: Tang, Bohan, et al.
Pubblicazione: (2025)
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
di: Wang, Jianghui, et al.
Pubblicazione: (2025)
di: Wang, Jianghui, et al.
Pubblicazione: (2025)
RELATE: A Schema-Agnostic Perceiver Encoder for Multimodal Relational Graphs
di: Meyer, Joe, et al.
Pubblicazione: (2025)
di: Meyer, Joe, et al.
Pubblicazione: (2025)
Optimal Singular Damage: Efficient LLM Inference in Low Storage Regimes
di: Alipour, Mohammadsajad, et al.
Pubblicazione: (2025)
di: Alipour, Mohammadsajad, et al.
Pubblicazione: (2025)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
di: Long, Xiang, et al.
Pubblicazione: (2026)
di: Long, Xiang, et al.
Pubblicazione: (2026)
GenoTEX: An LLM Agent Benchmark for Automated Gene Expression Data Analysis
di: Liu, Haoyang, et al.
Pubblicazione: (2024)
di: Liu, Haoyang, et al.
Pubblicazione: (2024)
BalanceBenchmark: A Survey for Multimodal Imbalance Learning
di: Xu, Shaoxuan, et al.
Pubblicazione: (2025)
di: Xu, Shaoxuan, et al.
Pubblicazione: (2025)
WAREX: Web Agent Reliability Evaluation on Existing Benchmarks
di: Kara, Su, et al.
Pubblicazione: (2025)
di: Kara, Su, et al.
Pubblicazione: (2025)
NaiAD: Initiate Data-Driven Research for LLM Advertising
di: Zhang, Yihang, et al.
Pubblicazione: (2026)
di: Zhang, Yihang, et al.
Pubblicazione: (2026)
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
di: Xu, Yang, et al.
Pubblicazione: (2026)
di: Xu, Yang, et al.
Pubblicazione: (2026)
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
di: Anurin, Andrey, et al.
Pubblicazione: (2024)
di: Anurin, Andrey, et al.
Pubblicazione: (2024)
Learning Multi-Agent Communication with Contrastive Learning
di: Lo, Yat Long, et al.
Pubblicazione: (2023)
di: Lo, Yat Long, et al.
Pubblicazione: (2023)
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
di: Vishwakarma, Harsh, et al.
Pubblicazione: (2025)
di: Vishwakarma, Harsh, et al.
Pubblicazione: (2025)
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
di: Rawles, Christopher, et al.
Pubblicazione: (2024)
di: Rawles, Christopher, et al.
Pubblicazione: (2024)
RouterBench: A Benchmark for Multi-LLM Routing System
di: Hu, Qitian Jason, et al.
Pubblicazione: (2024)
di: Hu, Qitian Jason, et al.
Pubblicazione: (2024)
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
di: Ma, Chang, et al.
Pubblicazione: (2024)
di: Ma, Chang, et al.
Pubblicazione: (2024)
Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
di: Bsharat, Sondos Mahmoud, et al.
Pubblicazione: (2025)
di: Bsharat, Sondos Mahmoud, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Survey on Evaluation of LLM-based Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2025) -
AUTOCT: Automating Interpretable Clinical Trial Prediction with LLM Agents
di: Liu, Fengze, et al.
Pubblicazione: (2025) -
TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks
di: Weng, Muyan, et al.
Pubblicazione: (2026) -
The Evaluation Game: Beyond Static LLM Benchmarking
di: Wang, Paul, et al.
Pubblicazione: (2026) -
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)