ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Jeon, YoungHoon, Kim, Suwan, Son, Haein, Lee, Sookbun, Jeong, Yeil, Lee, Unggi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Prediction to Application: Language Model-based Code Knowledge Tracing with Domain Adaptive Pre-Training and Automatic Feedback System with Pedagogical Prompting for Comprehensive Programming Education
di: Lee, Unggi, et al.
Pubblicazione: (2024)
di: Lee, Unggi, et al.
Pubblicazione: (2024)
GISclaw: A Comprehensive Open-Source LLM Agent System for Realistic Multi-Step Geospatial Analysis
di: Han, Jinzhen, et al.
Pubblicazione: (2026)
di: Han, Jinzhen, et al.
Pubblicazione: (2026)
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
di: Qiu, Jielin, et al.
Pubblicazione: (2025)
di: Qiu, Jielin, et al.
Pubblicazione: (2025)
Continuous Benchmark Generation for Evaluating Enterprise-scale LLM Agents
di: Saxena, Divyanshu, et al.
Pubblicazione: (2025)
di: Saxena, Divyanshu, et al.
Pubblicazione: (2025)
ClarEval: A Benchmark for Evaluating Clarification Skills of Code Agents under Ambiguous Instructions
di: Li, Jialin, et al.
Pubblicazione: (2026)
di: Li, Jialin, et al.
Pubblicazione: (2026)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
di: Garg, Spandan, et al.
Pubblicazione: (2025)
di: Garg, Spandan, et al.
Pubblicazione: (2025)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
di: Liu, Zhou, et al.
Pubblicazione: (2025)
di: Liu, Zhou, et al.
Pubblicazione: (2025)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
di: Tu, Xinming, et al.
Pubblicazione: (2026)
di: Tu, Xinming, et al.
Pubblicazione: (2026)
When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling
di: Islam, Niful, et al.
Pubblicazione: (2026)
di: Islam, Niful, et al.
Pubblicazione: (2026)
PyBench: Evaluating LLM Agent on various real-world coding tasks
di: Zhang, Yaolun, et al.
Pubblicazione: (2024)
di: Zhang, Yaolun, et al.
Pubblicazione: (2024)
DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
di: Xiao, Jingyu, et al.
Pubblicazione: (2025)
di: Xiao, Jingyu, et al.
Pubblicazione: (2025)
Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development
di: Zeng, Zhengran, et al.
Pubblicazione: (2025)
di: Zeng, Zhengran, et al.
Pubblicazione: (2025)
ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
di: Shen, Haiyang, et al.
Pubblicazione: (2024)
di: Shen, Haiyang, et al.
Pubblicazione: (2024)
PBT-Bench: Benchmarking AI Agents on Property-Based Testing
di: Jing, Lucas, et al.
Pubblicazione: (2026)
di: Jing, Lucas, et al.
Pubblicazione: (2026)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
di: Xiao, Yijia, et al.
Pubblicazione: (2025)
di: Xiao, Yijia, et al.
Pubblicazione: (2025)
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
di: Sun, Jiazheng, et al.
Pubblicazione: (2026)
di: Sun, Jiazheng, et al.
Pubblicazione: (2026)
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
di: He, Jiawei, et al.
Pubblicazione: (2026)
di: He, Jiawei, et al.
Pubblicazione: (2026)
CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
di: Lin, Haojia, et al.
Pubblicazione: (2025)
di: Lin, Haojia, et al.
Pubblicazione: (2025)
Evaluating LLM Agents on Automated Software Analysis Tasks
di: Bouzenia, Islem, et al.
Pubblicazione: (2026)
di: Bouzenia, Islem, et al.
Pubblicazione: (2026)
Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents
di: Xiang, Jiahong, et al.
Pubblicazione: (2026)
di: Xiang, Jiahong, et al.
Pubblicazione: (2026)
FireBench: Evaluating Instruction Following in Enterprise and API-Driven LLM Applications
di: Zhang, Yunfan, et al.
Pubblicazione: (2026)
di: Zhang, Yunfan, et al.
Pubblicazione: (2026)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
di: Zhang, Zehua, et al.
Pubblicazione: (2025)
di: Zhang, Zehua, et al.
Pubblicazione: (2025)
Learning to Ask: When LLM Agents Meet Unclear Instruction
di: Wang, Wenxuan, et al.
Pubblicazione: (2024)
di: Wang, Wenxuan, et al.
Pubblicazione: (2024)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
di: Gao, Zeyu, et al.
Pubblicazione: (2025)
di: Gao, Zeyu, et al.
Pubblicazione: (2025)
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
di: Merrill, Mike A., et al.
Pubblicazione: (2026)
di: Merrill, Mike A., et al.
Pubblicazione: (2026)
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
di: Guo, Hanyang, et al.
Pubblicazione: (2025)
di: Guo, Hanyang, et al.
Pubblicazione: (2025)
Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation
di: Pysklo, Hubert M., et al.
Pubblicazione: (2026)
di: Pysklo, Hubert M., et al.
Pubblicazione: (2026)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
di: Wang, Yubang, et al.
Pubblicazione: (2026)
di: Wang, Yubang, et al.
Pubblicazione: (2026)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
di: Duston, Titouan, et al.
Pubblicazione: (2025)
di: Duston, Titouan, et al.
Pubblicazione: (2025)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
di: Liu, Chenxu, et al.
Pubblicazione: (2026)
di: Liu, Chenxu, et al.
Pubblicazione: (2026)
A Self-Healing Framework for Reliable LLM-Based Autonomous Agents
di: Jeong, Cheonsu, et al.
Pubblicazione: (2026)
di: Jeong, Cheonsu, et al.
Pubblicazione: (2026)
From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection
di: Han, Junxiao, et al.
Pubblicazione: (2025)
di: Han, Junxiao, et al.
Pubblicazione: (2025)
HumanEvalComm: Benchmarking the Communication Competence of Code Generation for LLMs and LLM Agent
di: Wu, Jie JW, et al.
Pubblicazione: (2024)
di: Wu, Jie JW, et al.
Pubblicazione: (2024)
A Comprehensive Empirical Evaluation of Agent Frameworks on Code-centric Software Engineering Tasks
di: Yin, Zhuowen, et al.
Pubblicazione: (2025)
di: Yin, Zhuowen, et al.
Pubblicazione: (2025)
Environment-in-the-Loop: Rethinking Code Migration with LLM-based Agents
di: Li, Xiang, et al.
Pubblicazione: (2026)
di: Li, Xiang, et al.
Pubblicazione: (2026)
Managing Uncertainty in LLM-based Multi-Agent System Operation
di: Zhang, Man, et al.
Pubblicazione: (2026)
di: Zhang, Man, et al.
Pubblicazione: (2026)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent System
di: Hu, Li, et al.
Pubblicazione: (2025)
di: Hu, Li, et al.
Pubblicazione: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
di: Jiang, Hongchao, et al.
Pubblicazione: (2025)
di: Jiang, Hongchao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
From Prediction to Application: Language Model-based Code Knowledge Tracing with Domain Adaptive Pre-Training and Automatic Feedback System with Pedagogical Prompting for Comprehensive Programming Education
di: Lee, Unggi, et al.
Pubblicazione: (2024) -
GISclaw: A Comprehensive Open-Source LLM Agent System for Realistic Multi-Step Geospatial Analysis
di: Han, Jinzhen, et al.
Pubblicazione: (2026) -
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
di: Qiu, Jielin, et al.
Pubblicazione: (2025) -
Continuous Benchmark Generation for Evaluating Enterprise-scale LLM Agents
di: Saxena, Divyanshu, et al.
Pubblicazione: (2025) -
ClarEval: A Benchmark for Evaluating Clarification Skills of Code Agents under Ambiguous Instructions
di: Li, Jialin, et al.
Pubblicazione: (2026)