MuSACo: Multimodal Subject-Specific Selection and Adaptation for Expression Recognition with Co-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeeshan, Muhammad Osama, Gillet, Natacha, Koerich, Alessandro Lameiras, Pedersoli, Marco, Bremond, Francois, Granger, Eric
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912754746523648
author Zeeshan, Muhammad Osama
Gillet, Natacha
Koerich, Alessandro Lameiras
Pedersoli, Marco
Bremond, Francois
Granger, Eric
author_facet Zeeshan, Muhammad Osama
Gillet, Natacha
Koerich, Alessandro Lameiras
Pedersoli, Marco
Bremond, Francois
Granger, Eric
contents Personalized expression recognition (ER) involves adapting a machine learning model to subject-specific data for improved recognition of expressions with considerable interpersonal variability. Subject-specific ER can benefit significantly from multi-source domain adaptation (MSDA) methods, where each domain corresponds to a specific subject to improve model accuracy and robustness. Despite promising results, state-of-the-art MSDA approaches often overlook multimodal information or blend sources into a single domain, limiting subject diversity and failing to explicitly capture unique subject-specific characteristics. To address these limitations, we introduce MuSACo, a multimodal subject-specific selection and adaptation method for ER based on co-training. It leverages complementary information across multiple modalities and multiple source domains for subject-specific adaptation. This makes MuSACo particularly relevant for affective computing applications in digital health, such as patient-specific assessment for stress or pain, where subject-level nuances are crucial. MuSACo selects source subjects relevant to the target and generates pseudo-labels using the dominant modality for class-aware learning, in conjunction with a class-agnostic loss to learn from less confident target samples. Finally, source features from each modality are aligned, while only confident target features are combined. Experimental results on challenging multimodal ER datasets: BioVid, StressID, and BAH show that MuSACo outperforms UDA (blending) and state-of-the-art MSDA methods.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12522
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MuSACo: Multimodal Subject-Specific Selection and Adaptation for Expression Recognition with Co-Training
Zeeshan, Muhammad Osama
Gillet, Natacha
Koerich, Alessandro Lameiras
Pedersoli, Marco
Bremond, Francois
Granger, Eric
Computer Vision and Pattern Recognition
Personalized expression recognition (ER) involves adapting a machine learning model to subject-specific data for improved recognition of expressions with considerable interpersonal variability. Subject-specific ER can benefit significantly from multi-source domain adaptation (MSDA) methods, where each domain corresponds to a specific subject to improve model accuracy and robustness. Despite promising results, state-of-the-art MSDA approaches often overlook multimodal information or blend sources into a single domain, limiting subject diversity and failing to explicitly capture unique subject-specific characteristics. To address these limitations, we introduce MuSACo, a multimodal subject-specific selection and adaptation method for ER based on co-training. It leverages complementary information across multiple modalities and multiple source domains for subject-specific adaptation. This makes MuSACo particularly relevant for affective computing applications in digital health, such as patient-specific assessment for stress or pain, where subject-level nuances are crucial. MuSACo selects source subjects relevant to the target and generates pseudo-labels using the dominant modality for class-aware learning, in conjunction with a class-agnostic loss to learn from less confident target samples. Finally, source features from each modality are aligned, while only confident target features are combined. Experimental results on challenging multimodal ER datasets: BioVid, StressID, and BAH show that MuSACo outperforms UDA (blending) and state-of-the-art MSDA methods.
title MuSACo: Multimodal Subject-Specific Selection and Adaptation for Expression Recognition with Co-Training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.12522