Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bueno, Ivo, Hou, Ruikun, Bühler, Babette, Fütterer, Tim, Drimalla, James, Foster, Jonathan Kyle, Youngs, Peter, Gerjets, Peter, Trautwein, Ulrich, Kasneci, Enkelejda
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914194964611072
author Bueno, Ivo
Hou, Ruikun
Bühler, Babette
Fütterer, Tim
Drimalla, James
Foster, Jonathan Kyle
Youngs, Peter
Gerjets, Peter
Trautwein, Ulrich
Kasneci, Enkelejda
author_facet Bueno, Ivo
Hou, Ruikun
Bühler, Babette
Fütterer, Tim
Drimalla, James
Foster, Jonathan Kyle
Youngs, Peter
Gerjets, Peter
Trautwein, Ulrich
Kasneci, Enkelejda
contents Observation of classroom interactions can provide concrete feedback to teachers, but current methods rely on manual annotation, which is resource-intensive and hard to scale. This work explores AI-driven analysis of classroom recordings, focusing on multimodal instructional activity and discourse recognition as a foundation for actionable feedback. Using a densely annotated dataset of 164 hours of video and 68 lesson transcripts, we design parallel, modality-specific pipelines. For video, we evaluate zero-shot multimodal LLMs, fine-tuned vision-language models, and self-supervised video transformers on 24 activity labels. For transcripts, we fine-tune a transformer-based classifier with contextualized inputs and compare it against prompting-based LLMs on 19 discourse labels. To handle class imbalance and multi-label complexity, we apply per-label thresholding, context windows, and imbalance-aware loss functions. The results show that fine-tuned models consistently outperform prompting-based approaches, achieving macro-F1 scores of 0.577 for video and 0.460 for transcripts. These results demonstrate the feasibility of automated classroom analysis and establish a foundation for scalable teacher feedback systems.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00087
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data
Bueno, Ivo
Hou, Ruikun
Bühler, Babette
Fütterer, Tim
Drimalla, James
Foster, Jonathan Kyle
Youngs, Peter
Gerjets, Peter
Trautwein, Ulrich
Kasneci, Enkelejda
Computer Vision and Pattern Recognition
Observation of classroom interactions can provide concrete feedback to teachers, but current methods rely on manual annotation, which is resource-intensive and hard to scale. This work explores AI-driven analysis of classroom recordings, focusing on multimodal instructional activity and discourse recognition as a foundation for actionable feedback. Using a densely annotated dataset of 164 hours of video and 68 lesson transcripts, we design parallel, modality-specific pipelines. For video, we evaluate zero-shot multimodal LLMs, fine-tuned vision-language models, and self-supervised video transformers on 24 activity labels. For transcripts, we fine-tune a transformer-based classifier with contextualized inputs and compare it against prompting-based LLMs on 19 discourse labels. To handle class imbalance and multi-label complexity, we apply per-label thresholding, context windows, and imbalance-aware loss functions. The results show that fine-tuned models consistently outperform prompting-based approaches, achieving macro-F1 scores of 0.577 for video and 0.460 for transcripts. These results demonstrate the feasibility of automated classroom analysis and establish a foundation for scalable teacher feedback systems.
title Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.00087