Enhancing Ambiguous Dynamic Facial Expression Recognition with Soft Label-based Data Augmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kawamura, Ryosuke, Hayashi, Hideaki, Otake, Shunsuke, Takemura, Noriko, Nagahara, Hajime
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913912923881472
author Kawamura, Ryosuke
Hayashi, Hideaki
Otake, Shunsuke
Takemura, Noriko
Nagahara, Hajime
author_facet Kawamura, Ryosuke
Hayashi, Hideaki
Otake, Shunsuke
Takemura, Noriko
Nagahara, Hajime
contents Dynamic facial expression recognition (DFER) is a task that estimates emotions from facial expression video sequences. For practical applications, accurately recognizing ambiguous facial expressions -- frequently encountered in in-the-wild data -- is essential. In this study, we propose MIDAS, a data augmentation method designed to enhance DFER performance for ambiguous facial expression data using soft labels representing probabilities of multiple emotion classes. MIDAS augments training data by convexly combining pairs of video frames and their corresponding emotion class labels. This approach extends mixup to soft-labeled video data, offering a simple yet highly effective method for handling ambiguity in DFER. To evaluate MIDAS, we conducted experiments on both the DFEW dataset and FERV39k-Plus, a newly constructed dataset that assigns soft labels to an existing DFER dataset. The results demonstrate that models trained with MIDAS-augmented data achieve superior performance compared to the state-of-the-art method trained on the original dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20867
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Ambiguous Dynamic Facial Expression Recognition with Soft Label-based Data Augmentation
Kawamura, Ryosuke
Hayashi, Hideaki
Otake, Shunsuke
Takemura, Noriko
Nagahara, Hajime
Computer Vision and Pattern Recognition
Dynamic facial expression recognition (DFER) is a task that estimates emotions from facial expression video sequences. For practical applications, accurately recognizing ambiguous facial expressions -- frequently encountered in in-the-wild data -- is essential. In this study, we propose MIDAS, a data augmentation method designed to enhance DFER performance for ambiguous facial expression data using soft labels representing probabilities of multiple emotion classes. MIDAS augments training data by convexly combining pairs of video frames and their corresponding emotion class labels. This approach extends mixup to soft-labeled video data, offering a simple yet highly effective method for handling ambiguity in DFER. To evaluate MIDAS, we conducted experiments on both the DFEW dataset and FERV39k-Plus, a newly constructed dataset that assigns soft labels to an existing DFER dataset. The results demonstrate that models trained with MIDAS-augmented data achieve superior performance compared to the state-of-the-art method trained on the original dataset.
title Enhancing Ambiguous Dynamic Facial Expression Recognition with Soft Label-based Data Augmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.20867