Saved in:
Bibliographic Details
Main Authors: Bohy, Hugo, Tran, Minh, Haddad, Kevin El, Dutoit, Thierry, Soleymani, Mohammad
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.17502
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911119542583296
author Bohy, Hugo
Tran, Minh
Haddad, Kevin El
Dutoit, Thierry
Soleymani, Mohammad
author_facet Bohy, Hugo
Tran, Minh
Haddad, Kevin El
Dutoit, Thierry
Soleymani, Mohammad
contents Human social behaviors are inherently multimodal necessitating the development of powerful audiovisual models for their perception. In this paper, we present Social-MAE, our pre-trained audiovisual Masked Autoencoder based on an extended version of Contrastive Audio-Visual Masked Auto-Encoder (CAV-MAE), which is pre-trained on audiovisual social data. Specifically, we modify CAV-MAE to receive a larger number of frames as input and pre-train it on a large dataset of human social interaction (VoxCeleb2) in a self-supervised manner. We demonstrate the effectiveness of this model by finetuning and evaluating the model on different social and affective downstream tasks, namely, emotion recognition, laughter detection and apparent personality estimation. The model achieves state-of-the-art results on multimodal emotion recognition and laughter recognition and competitive results for apparent personality estimation, demonstrating the effectiveness of in-domain self-supervised pre-training. Code and model weight are available here https://github.com/HuBohy/SocialMAE.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17502
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice
Bohy, Hugo
Tran, Minh
Haddad, Kevin El
Dutoit, Thierry
Soleymani, Mohammad
Computer Vision and Pattern Recognition
Human social behaviors are inherently multimodal necessitating the development of powerful audiovisual models for their perception. In this paper, we present Social-MAE, our pre-trained audiovisual Masked Autoencoder based on an extended version of Contrastive Audio-Visual Masked Auto-Encoder (CAV-MAE), which is pre-trained on audiovisual social data. Specifically, we modify CAV-MAE to receive a larger number of frames as input and pre-train it on a large dataset of human social interaction (VoxCeleb2) in a self-supervised manner. We demonstrate the effectiveness of this model by finetuning and evaluating the model on different social and affective downstream tasks, namely, emotion recognition, laughter detection and apparent personality estimation. The model achieves state-of-the-art results on multimodal emotion recognition and laughter recognition and competitive results for apparent personality estimation, demonstrating the effectiveness of in-domain self-supervised pre-training. Code and model weight are available here https://github.com/HuBohy/SocialMAE.
title Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.17502