Saved in:
Bibliographic Details
Main Authors: Vedernikov, Alexander, Kumar, Puneet, Chen, Haoyu, Seppänen, Tapio, Li, Xiaobai
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.14749
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915625904898048
author Vedernikov, Alexander
Kumar, Puneet
Chen, Haoyu
Seppänen, Tapio
Li, Xiaobai
author_facet Vedernikov, Alexander
Kumar, Puneet
Chen, Haoyu
Seppänen, Tapio
Li, Xiaobai
contents Engagement recognition in video datasets, unlike traditional image classification tasks, is particularly challenged by subjective labels and noise limiting model performance. To overcome the challenges of subjective and noisy engagement labels, we propose a framework leveraging Vision Large Language Models (VLMs) to refine annotations and guide the training process. Our framework uses a questionnaire to extract behavioral cues and split data into high- and low-reliability subsets. We also introduce a training strategy combining curriculum learning with soft label refinement, gradually incorporating ambiguous samples while adjusting supervision to reflect uncertainty. We demonstrate that classical computer vision models trained on refined high-reliability subsets and enhanced with our curriculum strategy show improvements, highlighting benefits of addressing label subjectivity with VLMs. This method surpasses prior state of the art across engagement benchmarks such as EngageNet (three of six feature settings, maximum improvement of +1.21%), and DREAMS / PAFE with F1 gains of +0.22 / +0.06.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14749
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision Large Language Models Are Good Noise Handlers in Engagement Analysis
Vedernikov, Alexander
Kumar, Puneet
Chen, Haoyu
Seppänen, Tapio
Li, Xiaobai
Computer Vision and Pattern Recognition
Engagement recognition in video datasets, unlike traditional image classification tasks, is particularly challenged by subjective labels and noise limiting model performance. To overcome the challenges of subjective and noisy engagement labels, we propose a framework leveraging Vision Large Language Models (VLMs) to refine annotations and guide the training process. Our framework uses a questionnaire to extract behavioral cues and split data into high- and low-reliability subsets. We also introduce a training strategy combining curriculum learning with soft label refinement, gradually incorporating ambiguous samples while adjusting supervision to reflect uncertainty. We demonstrate that classical computer vision models trained on refined high-reliability subsets and enhanced with our curriculum strategy show improvements, highlighting benefits of addressing label subjectivity with VLMs. This method surpasses prior state of the art across engagement benchmarks such as EngageNet (three of six feature settings, maximum improvement of +1.21%), and DREAMS / PAFE with F1 gains of +0.22 / +0.06.
title Vision Large Language Models Are Good Noise Handlers in Engagement Analysis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.14749