Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chang, Andrew, Hu, Chenkai, Qi, Ji, Wei, Zhuojian, Zhang, Kexin, Akkaraju, Viswadruth, Poeppel, David, Freeman, Dustin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913997247217664
author Chang, Andrew
Hu, Chenkai
Qi, Ji
Wei, Zhuojian
Zhang, Kexin
Akkaraju, Viswadruth
Poeppel, David
Freeman, Dustin
author_facet Chang, Andrew
Hu, Chenkai
Qi, Ji
Wei, Zhuojian
Zhang, Kexin
Akkaraju, Viswadruth
Poeppel, David
Freeman, Dustin
contents Group conversations over videoconferencing are a complex social behavior. However, the subjective moments of negative experience, where the conversation loses fluidity or enjoyment remain understudied. These moments are infrequent in naturalistic data, and thus training a supervised learning (SL) model requires costly manual data annotation. We applied semi-supervised learning (SSL) to leverage targeted labeled and unlabeled clips for training multimodal (audio, facial, text) deep features to predict non-fluid or unenjoyable moments in holdout videoconference sessions. The modality-fused co-training SSL achieved an ROC-AUC of 0.9 and an F1 score of 0.6, outperforming SL models by up to 4% with the same amount of labeled data. Remarkably, the best SSL model with just 8% labeled data matched 96% of the SL model's full-data performance. This shows an annotation-efficient framework for modeling videoconference experience.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13971
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience
Chang, Andrew
Hu, Chenkai
Qi, Ji
Wei, Zhuojian
Zhang, Kexin
Akkaraju, Viswadruth
Poeppel, David
Freeman, Dustin
Audio and Speech Processing
Computation and Language
Human-Computer Interaction
Machine Learning
Multimedia
Group conversations over videoconferencing are a complex social behavior. However, the subjective moments of negative experience, where the conversation loses fluidity or enjoyment remain understudied. These moments are infrequent in naturalistic data, and thus training a supervised learning (SL) model requires costly manual data annotation. We applied semi-supervised learning (SSL) to leverage targeted labeled and unlabeled clips for training multimodal (audio, facial, text) deep features to predict non-fluid or unenjoyable moments in holdout videoconference sessions. The modality-fused co-training SSL achieved an ROC-AUC of 0.9 and an F1 score of 0.6, outperforming SL models by up to 4% with the same amount of labeled data. Remarkably, the best SSL model with just 8% labeled data matched 96% of the SL model's full-data performance. This shows an annotation-efficient framework for modeling videoconference experience.
title Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience
topic Audio and Speech Processing
Computation and Language
Human-Computer Interaction
Machine Learning
Multimedia
url https://arxiv.org/abs/2506.13971