Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zixin, Liu, Peng, Sheng, Rui, Li, Haobo, Tu, Jianhong, Deng, Xiaodong, Shum, Kashun, Liu, Dayiheng, Qu, Huamin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917517885177856
author Chen, Zixin
Liu, Peng
Sheng, Rui
Li, Haobo
Tu, Jianhong
Deng, Xiaodong
Shum, Kashun
Liu, Dayiheng
Qu, Huamin
author_facet Chen, Zixin
Liu, Peng
Sheng, Rui
Li, Haobo
Tu, Jianhong
Deng, Xiaodong
Shum, Kashun
Liu, Dayiheng
Qu, Huamin
contents Language agents are increasingly deployed in complex professional workflows, with tutoring emerging as a particularly high-stakes capability that remains largely unmeasured in existing benchmarks. Effective tutor agents require more than producing correct answers or executing accurate tool calls: a robust tutor must diagnose learner state, adapt support over time, make pedagogically justified decisions grounded in educational evidence, and execute interventions within realistic learning-management systems. We introduce EduAgentBench, a source-grounded benchmark for holistically evaluating tutor agents across the full scope of teaching work. It contains 150 quality-controlled tasks across three capability surfaces: professional pedagogical judgment, situated multi-turn tutoring, and Canvas-style teaching workflow completion. Tasks are constructed through a pedagogical-insight-driven pipeline and evaluated with complementary verification signals and human review. Across a comprehensive evaluation of frontier models, our findings reveal that current models are generally capable of bounded pedagogical judgment, but still fall short of professional teaching standards in situated tutoring and autonomous teaching-workflow execution. To our knowledge, EduAgentBench is the first theory-grounded and realistic benchmark for evaluating the holistic teaching capability of tutor agents, providing a measurement foundation for developing future tutor agents that can support realistic teaching work.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14322
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows
Chen, Zixin
Liu, Peng
Sheng, Rui
Li, Haobo
Tu, Jianhong
Deng, Xiaodong
Shum, Kashun
Liu, Dayiheng
Qu, Huamin
Artificial Intelligence
Language agents are increasingly deployed in complex professional workflows, with tutoring emerging as a particularly high-stakes capability that remains largely unmeasured in existing benchmarks. Effective tutor agents require more than producing correct answers or executing accurate tool calls: a robust tutor must diagnose learner state, adapt support over time, make pedagogically justified decisions grounded in educational evidence, and execute interventions within realistic learning-management systems. We introduce EduAgentBench, a source-grounded benchmark for holistically evaluating tutor agents across the full scope of teaching work. It contains 150 quality-controlled tasks across three capability surfaces: professional pedagogical judgment, situated multi-turn tutoring, and Canvas-style teaching workflow completion. Tasks are constructed through a pedagogical-insight-driven pipeline and evaluated with complementary verification signals and human review. Across a comprehensive evaluation of frontier models, our findings reveal that current models are generally capable of bounded pedagogical judgment, but still fall short of professional teaching standards in situated tutoring and autonomous teaching-workflow execution. To our knowledge, EduAgentBench is the first theory-grounded and realistic benchmark for evaluating the holistic teaching capability of tutor agents, providing a measurement foundation for developing future tutor agents that can support realistic teaching work.
title Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows
topic Artificial Intelligence
url https://arxiv.org/abs/2605.14322