M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nguyen-Phuoc, Long, Gaboriau, Renald, Delacroix, Dimitri, Navarro, Laurent
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910367246974976
author Nguyen-Phuoc, Long
Gaboriau, Renald
Delacroix, Dimitri
Navarro, Laurent
author_facet Nguyen-Phuoc, Long
Gaboriau, Renald
Delacroix, Dimitri
Navarro, Laurent
contents This paper introduces the M&M model, a novel multimodal-multitask learning framework, applied to the AVCAffe dataset for cognitive load assessment (CLA). M&M uniquely integrates audiovisual cues through a dual-pathway architecture, featuring specialized streams for audio and video inputs. A key innovation lies in its cross-modality multihead attention mechanism, fusing the different modalities for synchronized multitasking. Another notable feature is the model's three specialized branches, each tailored to a specific cognitive load label, enabling nuanced, task-specific analysis. While it shows modest performance compared to the AVCAffe's single-task baseline, M\&M demonstrates a promising framework for integrated multimodal processing. This work paves the way for future enhancements in multimodal-multitask learning systems, emphasizing the fusion of diverse data types for complex task handling.
format Preprint
id arxiv_https___arxiv_org_abs_2403_09451
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment
Nguyen-Phuoc, Long
Gaboriau, Renald
Delacroix, Dimitri
Navarro, Laurent
Computer Vision and Pattern Recognition
Multimedia
Sound
Audio and Speech Processing
This paper introduces the M&M model, a novel multimodal-multitask learning framework, applied to the AVCAffe dataset for cognitive load assessment (CLA). M&M uniquely integrates audiovisual cues through a dual-pathway architecture, featuring specialized streams for audio and video inputs. A key innovation lies in its cross-modality multihead attention mechanism, fusing the different modalities for synchronized multitasking. Another notable feature is the model's three specialized branches, each tailored to a specific cognitive load label, enabling nuanced, task-specific analysis. While it shows modest performance compared to the AVCAffe's single-task baseline, M\&M demonstrates a promising framework for integrated multimodal processing. This work paves the way for future enhancements in multimodal-multitask learning systems, emphasizing the fusion of diverse data types for complex task handling.
title M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment
topic Computer Vision and Pattern Recognition
Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2403.09451