ConfusionBench: An Expert-Validated Benchmark for Confusion Recognition and Localization in Educational Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Lu, Wang, Xiao, Frank, Mark, Setlur, Srirangaraj, Govindaraju, Venu, Nwogu, Ifeoma
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914404962926592
author Dong, Lu
Wang, Xiao
Frank, Mark
Setlur, Srirangaraj
Govindaraju, Venu
Nwogu, Ifeoma
author_facet Dong, Lu
Wang, Xiao
Frank, Mark
Setlur, Srirangaraj
Govindaraju, Venu
Nwogu, Ifeoma
contents Recognizing and localizing student confusion from video is an important yet challenging problem in educational AI. Existing confusion datasets suffer from noisy labels, coarse temporal annotations, and limited expert validation, which hinder reliable fine-grained recognition and temporally grounded analysis. To address these limitations, we propose a practical multi-stage filtering pipeline that integrates two stages of model-assisted screening, researcher curation, and expert validation to build a higher-quality benchmark for confusion understanding. Based on this pipeline, we introduce ConfusionBench, a new benchmark for educational videos consisting of a balanced confusion recognition dataset and a video localization dataset. We further provide zero-shot baseline evaluations of a representative open-source model and a proprietary model on clip-level confusion recognition, long-video confusion localization tasks. Experimental results show that the proprietary model performs better overall but tends to over-predict transitional segments, while the open-source model is more conservative and more prone to missed detections. In addition, the proposed student confusion report visualization can support educational experts in making intervention decisions and adapting learning plans accordingly. All datasets and related materials will be made publicly available on our project page.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17267
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ConfusionBench: An Expert-Validated Benchmark for Confusion Recognition and Localization in Educational Videos
Dong, Lu
Wang, Xiao
Frank, Mark
Setlur, Srirangaraj
Govindaraju, Venu
Nwogu, Ifeoma
Computer Vision and Pattern Recognition
Recognizing and localizing student confusion from video is an important yet challenging problem in educational AI. Existing confusion datasets suffer from noisy labels, coarse temporal annotations, and limited expert validation, which hinder reliable fine-grained recognition and temporally grounded analysis. To address these limitations, we propose a practical multi-stage filtering pipeline that integrates two stages of model-assisted screening, researcher curation, and expert validation to build a higher-quality benchmark for confusion understanding. Based on this pipeline, we introduce ConfusionBench, a new benchmark for educational videos consisting of a balanced confusion recognition dataset and a video localization dataset. We further provide zero-shot baseline evaluations of a representative open-source model and a proprietary model on clip-level confusion recognition, long-video confusion localization tasks. Experimental results show that the proprietary model performs better overall but tends to over-predict transitional segments, while the open-source model is more conservative and more prone to missed detections. In addition, the proposed student confusion report visualization can support educational experts in making intervention decisions and adapting learning plans accordingly. All datasets and related materials will be made publicly available on our project page.
title ConfusionBench: An Expert-Validated Benchmark for Confusion Recognition and Localization in Educational Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.17267