AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gioacchini, Luca, Siracusano, Giuseppe, Sanvito, Davide, Gashteovski, Kiril, Friede, David, Bifulco, Roberto, Lawrence, Carolin |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
MEDDxAgent: A Unified Modular Agent Framework for Explainable Automatic Differential Diagnosis
par: Rose, Daniel, et autres
Publié: (2025)
par: Rose, Daniel, et autres
Publié: (2025)
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
par: Gioacchini, Luca, et autres
Publié: (2024)
par: Gioacchini, Luca, et autres
Publié: (2024)
What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering
par: Errica, Federico, et autres
Publié: (2024)
par: Errica, Federico, et autres
Publié: (2024)
Compositional Steering of Large Language Models with Steering Tokens
par: Radevski, Gorjan, et autres
Publié: (2026)
par: Radevski, Gorjan, et autres
Publié: (2026)
Clean Up the Mess: Addressing Data Pollution in Cryptocurrency Abuse Reporting Services
par: Gomez, Gibran, et autres
Publié: (2024)
par: Gomez, Gibran, et autres
Publié: (2024)
TextMineX: Data, Evaluation Framework and Ontology-guided LLM Pipeline for Humanitarian Mine Action
par: Zhou, Chenyue, et autres
Publié: (2025)
par: Zhou, Chenyue, et autres
Publié: (2025)
Leveraging Open Information Extraction for More Robust Domain Transfer of Event Trigger Detection
par: Dukić, David, et autres
Publié: (2023)
par: Dukić, David, et autres
Publié: (2023)
LightPAL: Lightweight Passage Retrieval for Open Domain Multi-Document Summarization
par: Enomoto, Masafumi, et autres
Publié: (2024)
par: Enomoto, Masafumi, et autres
Publié: (2024)
Robust Text Classification: Analyzing Prototype-Based Networks
par: Sourati, Zhivar, et autres
Publié: (2023)
par: Sourati, Zhivar, et autres
Publié: (2023)
Evaluating Language Models as Synthetic Data Generators
par: Kim, Seungone, et autres
Publié: (2024)
par: Kim, Seungone, et autres
Publié: (2024)
Scaling Evaluation-time Compute with Reasoning Models as Evaluators
par: Kim, Seungone, et autres
Publié: (2025)
par: Kim, Seungone, et autres
Publié: (2025)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
par: Andriushchenko, Maksym, et autres
Publié: (2024)
par: Andriushchenko, Maksym, et autres
Publié: (2024)
Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design Decisions
par: Liu, Emmy, et autres
Publié: (2025)
par: Liu, Emmy, et autres
Publié: (2025)
AgentSquare: Automatic LLM Agent Search in Modular Design Space
par: Shang, Yu, et autres
Publié: (2024)
par: Shang, Yu, et autres
Publié: (2024)
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
par: Xu, Frank F., et autres
Publié: (2024)
par: Xu, Frank F., et autres
Publié: (2024)
Position: LLM Unlearning Benchmarks are Weak Measures of Progress
par: Thaker, Pratiksha, et autres
Publié: (2024)
par: Thaker, Pratiksha, et autres
Publié: (2024)
ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents
par: Guo, Zhihan, et autres
Publié: (2025)
par: Guo, Zhihan, et autres
Publié: (2025)
AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
par: Liu, Xuannan, et autres
Publié: (2026)
par: Liu, Xuannan, et autres
Publié: (2026)
Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
par: Wang, Siyuan, et autres
Publié: (2024)
par: Wang, Siyuan, et autres
Publié: (2024)
When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents
par: Qian, Lingfei, et autres
Publié: (2025)
par: Qian, Lingfei, et autres
Publié: (2025)
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
par: Shen, Yujiong, et autres
Publié: (2026)
par: Shen, Yujiong, et autres
Publié: (2026)
Conversational Health Agents: A Personalized LLM-Powered Agent Framework
par: Abbasian, Mahyar, et autres
Publié: (2023)
par: Abbasian, Mahyar, et autres
Publié: (2023)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
par: Fu, Lingyue, et autres
Publié: (2025)
par: Fu, Lingyue, et autres
Publié: (2025)
Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning
par: Cheng, Mingyue, et autres
Publié: (2025)
par: Cheng, Mingyue, et autres
Publié: (2025)
PABU: Progress-Aware Belief Update for Efficient LLM Agents
par: Jiang, Haitao, et autres
Publié: (2026)
par: Jiang, Haitao, et autres
Publié: (2026)
Robotouille: An Asynchronous Planning Benchmark for LLM Agents
par: Gonzalez-Pumariega, Gonzalo, et autres
Publié: (2025)
par: Gonzalez-Pumariega, Gonzalo, et autres
Publié: (2025)
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
par: Jang, Lawrence Keunho, et autres
Publié: (2026)
par: Jang, Lawrence Keunho, et autres
Publié: (2026)
Infusing Theory of Mind into Socially Intelligent LLM Agents
par: Hwang, EunJeong, et autres
Publié: (2025)
par: Hwang, EunJeong, et autres
Publié: (2025)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
par: Lee, Gyubok, et autres
Publié: (2025)
par: Lee, Gyubok, et autres
Publié: (2025)
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
par: Hu, Tiansheng, et autres
Publié: (2026)
par: Hu, Tiansheng, et autres
Publié: (2026)
Benchmark Test-Time Scaling of General LLM Agents
par: Li, Xiaochuan, et autres
Publié: (2026)
par: Li, Xiaochuan, et autres
Publié: (2026)
On Synthesizing Data for Context Attribution in Question Answering
par: Radevski, Gorjan, et autres
Publié: (2025)
par: Radevski, Gorjan, et autres
Publié: (2025)
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
par: Xi, Zhiheng, et autres
Publié: (2025)
par: Xi, Zhiheng, et autres
Publié: (2025)
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
par: Wang, Hanlin, et autres
Publié: (2025)
par: Wang, Hanlin, et autres
Publié: (2025)
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
par: Jeon, YoungHoon, et autres
Publié: (2026)
par: Jeon, YoungHoon, et autres
Publié: (2026)
Self-Improving LLM Agents at Test-Time
par: Acikgoz, Emre Can, et autres
Publié: (2025)
par: Acikgoz, Emre Can, et autres
Publié: (2025)
BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
par: Feng, Yunhao, et autres
Publié: (2026)
par: Feng, Yunhao, et autres
Publié: (2026)
AutoAgent: A Fully-Automated and Zero-Code Framework for LLM Agents
par: Tang, Jiabin, et autres
Publié: (2025)
par: Tang, Jiabin, et autres
Publié: (2025)
Improved LLM Agents for Financial Document Question Answering
par: Tan, Nelvin, et autres
Publié: (2025)
par: Tan, Nelvin, et autres
Publié: (2025)
MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning
par: Tang, Xiangru, et autres
Publié: (2025)
par: Tang, Xiangru, et autres
Publié: (2025)
Documents similaires
-
MEDDxAgent: A Unified Modular Agent Framework for Explainable Automatic Differential Diagnosis
par: Rose, Daniel, et autres
Publié: (2025) -
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
par: Gioacchini, Luca, et autres
Publié: (2024) -
What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering
par: Errica, Federico, et autres
Publié: (2024) -
Compositional Steering of Large Language Models with Steering Tokens
par: Radevski, Gorjan, et autres
Publié: (2026) -
Clean Up the Mess: Addressing Data Pollution in Cryptocurrency Abuse Reporting Services
par: Gomez, Gibran, et autres
Publié: (2024)