M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866910367246974976 |
|---|---|
| author | Nguyen-Phuoc, Long Gaboriau, Renald Delacroix, Dimitri Navarro, Laurent |
| author_facet | Nguyen-Phuoc, Long Gaboriau, Renald Delacroix, Dimitri Navarro, Laurent |
| contents | This paper introduces the M&M model, a novel multimodal-multitask learning framework, applied to the AVCAffe dataset for cognitive load assessment (CLA). M&M uniquely integrates audiovisual cues through a dual-pathway architecture, featuring specialized streams for audio and video inputs. A key innovation lies in its cross-modality multihead attention mechanism, fusing the different modalities for synchronized multitasking. Another notable feature is the model's three specialized branches, each tailored to a specific cognitive load label, enabling nuanced, task-specific analysis. While it shows modest performance compared to the AVCAffe's single-task baseline, M\&M demonstrates a promising framework for integrated multimodal processing. This work paves the way for future enhancements in multimodal-multitask learning systems, emphasizing the fusion of diverse data types for complex task handling. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_09451 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment Nguyen-Phuoc, Long Gaboriau, Renald Delacroix, Dimitri Navarro, Laurent Computer Vision and Pattern Recognition Multimedia Sound Audio and Speech Processing This paper introduces the M&M model, a novel multimodal-multitask learning framework, applied to the AVCAffe dataset for cognitive load assessment (CLA). M&M uniquely integrates audiovisual cues through a dual-pathway architecture, featuring specialized streams for audio and video inputs. A key innovation lies in its cross-modality multihead attention mechanism, fusing the different modalities for synchronized multitasking. Another notable feature is the model's three specialized branches, each tailored to a specific cognitive load label, enabling nuanced, task-specific analysis. While it shows modest performance compared to the AVCAffe's single-task baseline, M\&M demonstrates a promising framework for integrated multimodal processing. This work paves the way for future enhancements in multimodal-multitask learning systems, emphasizing the fusion of diverse data types for complex task handling. |
| title | M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment |
| topic | Computer Vision and Pattern Recognition Multimedia Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2403.09451 |