Audio-Visual Compound Expression Recognition Method based on Late Modality Fusion and Rule-based Decision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ryumina, Elena, Markitantov, Maxim, Ryumin, Dmitry, Kaya, Heysem, Karpov, Alexey
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913290306715648
author Ryumina, Elena
Markitantov, Maxim
Ryumin, Dmitry
Kaya, Heysem
Karpov, Alexey
author_facet Ryumina, Elena
Markitantov, Maxim
Ryumin, Dmitry
Kaya, Heysem
Karpov, Alexey
contents This paper presents the results of the SUN team for the Compound Expressions Recognition Challenge of the 6th ABAW Competition. We propose a novel audio-visual method for compound expression recognition. Our method relies on emotion recognition models that fuse modalities at the emotion probability level, while decisions regarding the prediction of compound expressions are based on predefined rules. Notably, our method does not use any training data specific to the target task. Thus, the problem is a zero-shot classification task. The method is evaluated in multi-corpus training and cross-corpus validation setups. Using our proposed method is achieved an F1-score value equals to 22.01% on the C-EXPR-DB test subset. Our findings from the challenge demonstrate that the proposed method can potentially form a basis for developing intelligent tools for annotating audio-visual data in the context of human's basic and compound emotions.
format Preprint
id arxiv_https___arxiv_org_abs_2403_12687
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Audio-Visual Compound Expression Recognition Method based on Late Modality Fusion and Rule-based Decision
Ryumina, Elena
Markitantov, Maxim
Ryumin, Dmitry
Kaya, Heysem
Karpov, Alexey
Computer Vision and Pattern Recognition
Machine Learning
This paper presents the results of the SUN team for the Compound Expressions Recognition Challenge of the 6th ABAW Competition. We propose a novel audio-visual method for compound expression recognition. Our method relies on emotion recognition models that fuse modalities at the emotion probability level, while decisions regarding the prediction of compound expressions are based on predefined rules. Notably, our method does not use any training data specific to the target task. Thus, the problem is a zero-shot classification task. The method is evaluated in multi-corpus training and cross-corpus validation setups. Using our proposed method is achieved an F1-score value equals to 22.01% on the C-EXPR-DB test subset. Our findings from the challenge demonstrate that the proposed method can potentially form a basis for developing intelligent tools for annotating audio-visual data in the context of human's basic and compound emotions.
title Audio-Visual Compound Expression Recognition Method based on Late Modality Fusion and Rule-based Decision
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2403.12687