Learning in Focus: Detecting Behavioral and Collaborative Engagement Using Vision Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Penchala, Sindhuja, Kontham, Saketh Reddy, Bhattacharjee, Prachi, Mahmoodi, Nima, Fonseca, Daniel, Karami, Sareh, Ghahremani, Mehdi, Perkins, Andy D., Rahimi, Shahram, Golilarz, Noorbakhsh Amiri
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918248650375168
author Penchala, Sindhuja
Kontham, Saketh Reddy
Bhattacharjee, Prachi
Mahmoodi, Nima
Fonseca, Daniel
Karami, Sareh
Ghahremani, Mehdi
Perkins, Andy D.
Rahimi, Shahram
Golilarz, Noorbakhsh Amiri
author_facet Penchala, Sindhuja
Kontham, Saketh Reddy
Bhattacharjee, Prachi
Mahmoodi, Nima
Fonseca, Daniel
Karami, Sareh
Ghahremani, Mehdi
Perkins, Andy D.
Rahimi, Shahram
Golilarz, Noorbakhsh Amiri
contents In early childhood education, accurately detecting collaborative and behavioral engagement is essential to foster meaningful learning experiences. This paper presents an AI driven approach that leverages Vision Transformers (ViTs) to automatically classify children s engagement using visual cues such as gaze direction, interaction, and peer collaboration. Utilizing the ChildPlay gaze dataset, our method is trained on annotated video segments to classify behavioral and collaborative engagement states (e.g., engaged, not engaged, collaborative, not collaborative). We evaluated six state of the art transformer models: Vision Transformer (ViT), Data efficient Image Transformer (DeiT), Swin Transformer, VitGaze, APVit and GazeTR. Among these, the Swin Transformer achieved the highest classification performance with an accuracy of 97.58 percent, demonstrating its effectiveness in modeling local and global attention. Our results highlight the potential of transformer based architectures for scalable, automated engagement analysis in real world educational settings.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15782
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning in Focus: Detecting Behavioral and Collaborative Engagement Using Vision Transformers
Penchala, Sindhuja
Kontham, Saketh Reddy
Bhattacharjee, Prachi
Mahmoodi, Nima
Fonseca, Daniel
Karami, Sareh
Ghahremani, Mehdi
Perkins, Andy D.
Rahimi, Shahram
Golilarz, Noorbakhsh Amiri
Neurons and Cognition
Artificial Intelligence
In early childhood education, accurately detecting collaborative and behavioral engagement is essential to foster meaningful learning experiences. This paper presents an AI driven approach that leverages Vision Transformers (ViTs) to automatically classify children s engagement using visual cues such as gaze direction, interaction, and peer collaboration. Utilizing the ChildPlay gaze dataset, our method is trained on annotated video segments to classify behavioral and collaborative engagement states (e.g., engaged, not engaged, collaborative, not collaborative). We evaluated six state of the art transformer models: Vision Transformer (ViT), Data efficient Image Transformer (DeiT), Swin Transformer, VitGaze, APVit and GazeTR. Among these, the Swin Transformer achieved the highest classification performance with an accuracy of 97.58 percent, demonstrating its effectiveness in modeling local and global attention. Our results highlight the potential of transformer based architectures for scalable, automated engagement analysis in real world educational settings.
title Learning in Focus: Detecting Behavioral and Collaborative Engagement Using Vision Transformers
topic Neurons and Cognition
Artificial Intelligence
url https://arxiv.org/abs/2508.15782